Why AI pilots stall between the demo and production
The gap is rarely the model. It is grounding, evaluation and ownership — three things a pilot can skip and a production system cannot.
A pilot that impresses a steering committee and a system a customer can touch are different artifacts. The distance between them is predictable, and it is almost never the model.
Grounding is the first wall
A demo answers from a curated document set. Production answers from the real estate: systems that disagree with each other, documents nobody has owned in three years, and access rules that differ per user. Retrieval quality collapses on contact with that, and the fix is data work rather than prompt work.
This is why the data practice and the AI practice are the same conversation. An agent pointed at an ungoverned estate produces confident answers from whichever copy it happened to retrieve.
Evaluation is the second
Pilots are evaluated by demo. Production needs a test set, a scoring method, and a number that has to hold before a change ships. Without it there is no way to tell whether a prompt change improved the system or moved the failure somewhere nobody looked.
The test set does not need to be large. It needs to contain the cases that would embarrass you.
Ownership is the third, and the one that actually kills projects
Pilots are owned by whoever was curious. Production needs someone whose job includes it on a Tuesday in eight months. When a pilot stalls and everyone agrees it was promising, this is usually the missing piece rather than any technical one.
Before the next pilot, it is worth asking who would own it if it worked. If the answer is unclear, that is useful to know at the start rather than at the end.
