Why AI Pilots Stall After the First Success
The pilot worked, and four months later nothing has moved. Three reasons that is nearly always true, and the boring document that unsticks it.
The pilot worked. That is what makes this hard to talk about.
Someone in the firm took one job — quoting, chasing records, writing up notes, triaging the inbox — pointed a model at it, and it came back faster and better than expected. Everyone in the room agreed it was impressive. A decision was made to roll it out more widely.
Then, four months on, the same tool is being used by the same one person for the same one job, and nobody can quite say what happened.
This is the most common shape of AI failure in a UK SME. Not a project that blows up. A project that succeeds once and then does not move.
The pilot succeeded because it was small
A pilot works partly because of the tool and mostly because of the conditions around it. One motivated person, one well-understood task, no integration, no handover, and a very forgiving definition of "done" — because everyone knows it is an experiment.
None of those conditions survive contact with the second use case. The second one goes to someone who did not volunteer, on a task nobody has written down, with an output that has to be right rather than interesting.
So the pilot did not really prove that the tool works in your firm. It proved that the tool works for the person who wanted it to. Those are different findings, and treating the first as the second is what produces the stall.
The three things that actually stop it
Nobody owns the second one. The pilot had a champion — someone who cared enough to sit with it through the bad first week. Rollout has a mailing list. The task that made the pilot succeed, which is patiently reworking prompts and inputs until the output is reliable, has no owner the second time round, and it does not happen by itself.
The next process was never written down. The pilot task was legible: everyone knew what a good quote looked like. The next one along is usually a job that lives in someone's head, done slightly differently by four people, with exceptions nobody has articulated. You cannot automate a process you cannot describe, and discovering that is a management problem wearing a technology costume.
Nothing was ever measured, so nothing can be justified. Ask what the pilot saved and you will usually get "loads of time" or a number someone estimated in a meeting. That is enough to authorise an experiment and nowhere near enough to authorise a budget line, a subscription for thirty people, or the hours it takes to change how a team works. Without a real before-and-after, the second phase has to be argued from enthusiasm, and enthusiasm runs out around the same time the novelty does.
The stall is usually about the inventory
Underneath all three is a more boring problem: most firms cannot say what they already run.
By the time a pilot has succeeded, staff have typically been using AI privately for months — on personal accounts, on consumer tiers, on data nobody approved. The official pilot lands on top of an unofficial estate that nobody has mapped. Every attempt to expand runs into a version of "wait, are we already doing this somewhere?" and quietly loses a fortnight.
Writing the inventory down is unglamorous and takes an afternoon. It is also the single change that most reliably unsticks a stalled programme, because it converts a vague expansion into a specific list of decisions. What we use, who uses it, what data goes in, what we are replacing. AI governance for SMEs covers what that document contains.
What to do instead of a bigger pilot
The instinct after a stall is to run a more ambitious trial. That usually fails in the same way, more expensively.
The alternative is smaller and duller. Pick the second task on the basis of whether it is written down, not whether it is exciting. Give it a named owner with real hours in their week, not a mention in someone's objectives. Measure one number before you start — minutes per job, error rate, days to turn something round — and measure the same number six weeks later, even if the answer is unflattering.
And accept the finding when it comes. Some processes are not worth automating, and knowing which ones is most of the value of doing this properly.
The pattern behind it
A pilot is a test of a tool at a fixed moment. A rollout is a commitment to a system that keeps changing underneath you — new model versions, new behaviour, new failure modes, on somebody else's release schedule. The controls and assumptions you set during the pilot start ageing the day it ends. The board-level version of that argument is in static controls and live models.
If your programme has gone quiet since its first win, the useful question is not which tool to try next. It is which of the three gaps above you actually have. The AI readiness scorecard takes about five minutes, scores you on ownership, process clarity and measurement, and shows the result straight away. No email required.
Ready to integrate AI into your business?
See how Model Context Protocol (MCP) can connect your AI assistant to all your business tools. Book a call with our team to discuss your specific needs.
Book a Call (opens in a new tab)