The pilot worked and the rollout did not
A pilot succeeds because everyone involved knows it is a pilot. The properties that make it work — a tolerant audience, a small corpus, manual fixes at the edges — are exactly the ones the rollout cannot inherit.
The gap nobody plans for
The pilot finishes well. A group of volunteers used the system for a quarter and reported that it saved them time. The decision is made to roll it out to everyone. Within a month, usage is a fraction of what the pilot predicted and the pilot group is defending a tool that most of the organisation has quietly abandoned.
The system did not change. The population did, and the population was never part of the design.
Three properties a pilot cannot tell you about
Volunteers are tolerant. People who opt into a pilot are interested in it working, and they compensate for failures by rephrasing, retrying and checking results. Nothing in their behaviour tells you whether the same quality bar survives a user who was told to use the tool and has no interest in its success.
Manual fixes are invisible requirements. During a pilot, somebody re-uploads a document, corrects a prompt, or fixes an index entry by hand every week. Those interventions keep the measured quality high while hiding the engineering work the rollout will require. The honest inventory is a list of what humans did around the system, and each item is either automated or accepted as permanent cost.
Adoption is measured wrong. Satisfaction surveys measure enthusiasm, which pilots have in surplus. What predicts rollout is repetition: the share of users who come back unprompted in consecutive weeks, and the share of the target workflow that actually moves into the tool rather than being duplicated around it.
What to collect before the rollout decision
- Weekly returning users, not total users. A pilot with thirty participants and five who return is a five-person product with enthusiastic spectators.
- A written list of every manual intervention performed during the pilot, with an owner and a decision for each.
- The failure rate on the users least like the pilot group — the ones who joined late, or who were added because of a reorganisation rather than because they asked.
- The cost of the pilot at rollout scale, including the support load that currently lands on the pilot team as a side effect of their enthusiasm.
- One workflow, defined narrowly enough that success is unambiguous, instead of a platform that is expected to absorb several.
Rolling out anyway
The rollout is still usually the right decision; the mistake is treating it as an expansion of the same thing rather than a new deployment with a new population. Run it in cohorts, keep the instrumentation that measured returning users during the pilot, and staff the first cohort with the same attention the pilot received. The failure mode is not a bad product, it is a good pilot mistaken for evidence about strangers.
What to do about it
- Pilots succeed partly because participants tolerate a lower quality bar
- Any manual fix applied during the pilot is an unbuilt requirement
- Measure adoption during the pilot, not satisfaction