Why Your AI Pilot Died After the Demo (And What a Production Agent Needs)
The pattern
A vendor demo, or an internal build, that did something genuinely impressive. A pilot approved. Seats bought. A kickoff meeting. Then a slow fade: a few people tried it, the results were uneven, nobody was sure what it was allowed to change, and the seats are still being paid for.
We hear this story on most first calls. The team usually blames the model. In our experience, the model is the last thing that went wrong.
Here are the five places a pilot actually dies, in the order they usually happen.
1. It read from data nobody trusted
The demo ran on clean sample data. The pilot ran on your CRM. Those are different things.
If reps don't trust the account list, they won't trust a summary generated from it. A pilot that reads from an untrusted record inherits the distrust, and nothing the model does can fix that. The fix is upstream. There are six things a CRM needs before an agent touches it, and most stalled pilots fail at least three.
2. It wasn't allowed to write anything
Most pilots are read-only by design. The agent drafts, suggests, summarizes. A person copies the result somewhere.
That's a sensible way to start and a terrible place to stop. If the output never lands in the system of record, the agent creates work instead of removing it. Reps stop copying. Usage drops. The dashboard says "adoption problem." It's a wiring problem.
A production agent writes. That means the build has to answer what it writes, where, and what stops a bad write. Every agent we've built has a confidence score and a threshold. Above it, the record is written and logged. Below it, a person reviews. The wellness-plan platform we built writes: intake and lab work go in, and a compliant plan lands on the client record in about a minute, after the scrubber passes it. Nobody copies anything.
3. Nobody could see what it decided
An agent makes a call, a person disagrees, and there's no way to see what the agent saw. So the person concludes it's random. After a few of those, they stop using it.
Production agents keep a log: the input, the fields extracted, the confidence, the decision, the write. When a rep disagrees, you can look. Usually one of two things is true. The input really was ambiguous, and the threshold should catch it next time. Or the instruction was wrong, and you fix the instruction. Either way it gets better, and the rep can see it get better.
A pilot with no log can't improve. It can only be abandoned.
4. It had no owner after launch
The vendor's team ran the demo. The internal champion ran the pilot on top of their real job. When the champion got busy, nobody was watching the agent, and an agent nobody watches drifts.
Prompts don't change on their own, but everything around them does. A picklist value gets renamed. A new product line appears that the instruction never mentioned. A phone provider changes its transcript format. Each one degrades the output a little, and with no owner, nobody notices until it's bad enough to switch off.
This is why our AI Implementation engagements end with a choice: a full handoff with docs and rollback, or a flat monthly plan where we host, monitor, and keep improving the agent. The support bot hardening we run for a battery manufacturer shows what watching looks like: a regression set of questions re-asked every week. Dead links in the bot's answers went from 77 percent to 3 percent in three weeks, and the weekly report is what keeps them there.
5. It was scoped as "AI" instead of as a job
"Pilot AI in sales" isn't a scope. "Turn every lab PDF a clinic uploads into named markers on the client record, within a minute, at an accuracy target you set, with anything uncertain queued for review" is a scope. The first one can't fail because it can't succeed. The second one ships or it doesn't.
Pilots that start with the technology and look for a use sprawl. Pilots that start with one job, one input, one write, and one number tend to finish. The use case library is a list of jobs sized that way, because that's the size that ships.
What a production agent needs
The six CRM checks above cover the record. On top of those, a production agent needs three things a pilot almost never has:
- A visible decision trail. Input, extraction, confidence, decision. Per record, so a rep can check the agent's work instead of guessing at it.
- An owner after launch. Yours or ours. Someone reading the weekly numbers and changing the instruction when the business changes.
- A job-sized scope. One input, one write, one number that says whether it worked.
None of that is glamorous. All of it is the difference between a demo and a system your team stops noticing because it just runs.
Where to start
If you have AI seats nobody turned on, or a pilot that's been "almost ready" for two quarters, the fastest path is usually not a new pilot. It's a read-only audit of what the last one was reading from and writing to. That tells you which of the five killed it, and whether it's worth reviving.
Book a 30-minute call and bring the pilot. We'll tell you where it stopped, and if the honest answer is that you don't need AI for this job, we'll say that instead.
Get Salesforce insights in your inbox.
Practical tips on org health, implementation, and CRM strategy. No spam. Unsubscribe anytime.