A pilot nobody is allowed to stop is not an experiment
The recognisable pattern in an established business: a pilot is approved, runs for a quarter, produces mixed results, and rolls out anyway. The success criteria were written after the data arrived. The sponsor’s standing is attached to the outcome. Nobody in the room has the authority, or the incentive, to call it.
Two things convert a pilot into an experiment, and both are organisational rather than technical. First, a kill criterion agreed and recorded before the work starts, stated as a number or an observable behaviour. Second, a named person who may invoke it, whose position does not suffer for doing so.
Test the assumption that would kill the idea, not the one that is easiest to test
Ideas rest on a stack of assumptions, and teams naturally test the ones nearest their own competence. Engineers test feasibility. Marketers test messaging. Meanwhile the assumption carrying the actual risk goes untouched, usually because it belongs to someone else’s domain.
Rank assumptions by uncertainty multiplied by consequence, and test the top one first. In an established business the top item is rarely whether the thing can be built. It is more often whether anyone will change their working routine to use it, whether operations can absorb the exception volume it creates, or whether the commercial terms survive contact with procurement.
The smallest useful test is usually not software
The instinct is to build a thin version of the product. Frequently the better instrument is a person doing the job by hand under the proposed rules. You are testing whether the outcome has value, not whether it can be automated, and those are separate hypotheses with separate risks.
A worked example. Before building an automated quoting tool, have someone produce quotes manually for a month using exactly the rules the tool would apply, and count how often a human had to override them. That override rate determines whether automation is viable at all. Build first and you discover it after the money is spent, at which point the tool is quietly abandoned in favour of a spreadsheet.
This is also the honest way to scope AI and automation work: measure the exception rate before committing to the model, because exceptions are where the economics live.
Measure behaviour that cost somebody something
Enthusiasm is free, which is exactly why it is not evidence. Positive feedback in a review session tells you the idea is socially acceptable, not that it is valuable. The signals worth counting are the ones that cost the person something they could not easily give away.
Look for a changed routine, a committed budget line, a recurring diary slot, a signature, a process step someone stopped doing because the new thing replaced it. In an internal context the strongest evidence is often a team refusing to go back to the old method when the pilot pauses: the complaint is the data.
The counter-argument: lean is a poor fit for infrastructure
This is the pushback the method usually escapes. You cannot run a minimum viable general ledger. You cannot iterate towards a correct identity platform, or discover an ERP migration through progressive experimentation. Where a system must be complete to be correct, and where reversal is expensive or impossible, iterative discovery is the wrong instrument.
The distinction is between demand-side uncertainty and build-side uncertainty. Lean is a method for resolving uncertainty about what people will do: will they buy it, adopt it, pay for it. It is not a method for resolving uncertainty about whether a complex, tightly coupled, correctness-critical system will work. That is answered by design rigour, rehearsal, staged cutover and testing, which look like the opposite of lean and are.
Budget by reversal cost, not by size
The useful frame is not how much a decision costs but how much it costs to unmake. Reversible decisions (a pricing rule, a landing page, a workflow, a supplier trial) should be made quickly, at low ceremony, by the people closest to them. Irreversible ones (a system of record, a data model, an entity structure, a long contract) deserve slow, expensive, senior deliberation.
Most organisations get this backwards in a consistent direction. They subject small reversible choices to committee, which is where lead time goes, and take large irreversible ones under time pressure at the end of a selection process, which is where money goes.
The transition from proving to scaling kills more ideas than the experiment does
A pilot runs on goodwill, workarounds and a sponsor’s attention. Scaling requires none of those and all of the things a pilot deliberately skipped: an owner, support arrangements, access control, data quality that holds at volume, exception handling, and a security review. The step change is not in the technology; it is in everything around it.
Build the list of what changes at scale at the point the pilot begins, not when it succeeds. It makes the pilot more honest, because some ideas will be visibly unscalable before any money is spent, and it removes the awkward moment where a celebrated pilot dies in an architecture review.
Keep the negative results
Businesses record what worked and forget what did not, so the same idea returns every two or three years with a new sponsor and no memory of why it was stopped. The cost is invisible, recurring, and entirely avoidable.
Keep a short record for each experiment: the assumption tested, the criterion set in advance, the result, and the reason for stopping. A paragraph is enough. Its value appears the first time someone proposes an idea that was already tested, and either they learn something or they explain what has changed since, both of which are better than starting again.
This is also the raw material for improving the method itself. Reviewing a year of kill decisions tells you whether your criteria were too loose, too tight, or written to be unfalsifiable. Most organisations discover the third. Consulting engagements that leave behind this habit are worth more than the individual decisions they informed.
Reduce the cost of being wrong
Link-IT helps leadership teams design experiments with real kill criteria, and knows where the method stops applying, so foundational work gets the rigour it needs instead.
Discuss your next stage