Why your AI pilot never reached production
The demo worked. Everyone in the room agreed it was impressive. Nine months later it is still a demo. This is the most common shape of an AI project we are asked to rescue, and the cause is almost never the model.
Across the assessments we have run, stalled pilots fail on one of four things. None of them is about capability. All of them are decided before a line of code is written.
1. There is no evaluation set
A demo is judged by whether the room nods. Production is judged by whether it is right often enough to trust, and nobody can answer that without a set of cases with known answers.
The work is unglamorous: sit with the people who do the task today, collect one or two hundred real examples, and write down what a correct response looks like for each. That set becomes the thing you run against every change. Without it you cannot tell an improvement from a regression, which means you cannot safely change anything, which means the pilot freezes.
If you take one thing from this article: build the evaluation set before you choose a model. It will outlive whichever model you pick.
2. Nobody owns the output
Ask who is accountable when the assistant is wrong. In a stalled project the answer is a shrug, or worse, "the AI". Systems without an owner do not get deployed, because no one is willing to sign for them.
The fix is organisational, not technical. Name a person in the business — not in IT — who owns the quality of the output, give them the evaluation results, and give them the authority to turn it off. Adoption follows accountability far more reliably than it follows accuracy.
3. There is no path to the data
Pilots are built on an export. Somebody pulls a spreadsheet, the assistant answers beautifully against it, and then the question arrives: how does this stay current?
That question is usually harder than the AI work. It means agreeing a source of truth, getting read access through the right boundary, handling permissions so the assistant cannot surface something a user should not see, and keeping it fresh on a schedule someone maintains. Teams that treat this as an afterthought lose months.
- Which system is authoritative for this data, and who says so?
- How does access work without copying the data somewhere less governed?
- If a document is restricted, does the assistant respect that restriction?
- What is the acceptable staleness — minutes, hours, or a day?
4. Cost is unbounded
A pilot serving five people costs nothing worth measuring. The same design serving five hundred, with long documents in every request, can produce a bill that ends the project in its second month.
Model the cost per interaction before you build, then set a hard ceiling in the system itself — a budget per user, per day, per workflow, enforced in code rather than watched on a dashboard. Caching repeated context and routing simple requests to smaller models routinely cuts spend by more than half without a change anyone notices.
How we de-risk it
Our discovery engagement answers all four questions before committing to a build. Two weeks, fixed fee: we map the workflow, collect the evaluation set with the people who do the work, establish the data path and its permission model, and model cost at realistic volume.
Sometimes the honest answer is that the process should be fixed rather than automated. That finding is worth more than a pilot that never ships, and you get it in a fortnight rather than a year.
Related practice
AI & Automation
Put AI to work on the processes that actually cost you money.
What this involvesRead next
Working on something this touches? We start with a two-week, fixed-fee discovery.
Talk to an engineer