← All posts
17 August 2026Gustforward Marketing Team

The pilot that worked and the rollout that didn't

Your AI pilot succeeded and then nothing happened. That's not a technology failure — it's four organisational ones, and they're predictable.

A yellow path on a dark navy field that widens, then frays into thin scattered threads

There's a specific kind of AI failure that doesn't look like failure.

The pilot goes well. The demo lands. Everyone agrees it's promising. A steering group is formed. And then, six months later, the thing is still running for the same four enthusiasts who ran the pilot, and nobody can quite explain what happened.

Nothing broke. That's what makes it hard to diagnose. The technology worked exactly as advertised and the value never arrived.

Four reasons, in roughly the order they bite.

1. The pilot ran on the easy half of the work

Pilots get chosen for pilot-ability. Clean data, cooperative team, a workflow someone already understands well enough to explain in a meeting. That's sensible — you're testing feasibility.

But feasibility on the tidy half tells you almost nothing about the messy half, and the messy half is most of the business. The branch that never adopted the CRM. The client whose contract has three bespoke exceptions. The team that does the same job with different vocabulary.

The tell: your pilot's success metric was "did it work," when it should have been "what fraction of the real distribution did we just cover, and what does the rest look like?"

The fix: during the pilot, deliberately run the agent in observe mode over the ugly cases too. You don't have to handle them yet. You have to count them, before you promise a rollout timeline based on the tidy sample.

2. Nobody owned it after the pilot team left

A pilot has an owner by definition — the person whose idea it was. Rollout often has a committee, which is a different thing.

Agents need an operator. Someone who reads the escalation queue, notices the approval rate dipping, decides whether the new edge case deserves a rule or a shrug. This is a real job, maybe a few hours a week, and it is nobody's job by default.

Where it goes wrong is subtle: the agent doesn't break, it drifts. A new service line launches and the agent doesn't know about it. A pricing rule changes and nowhere in the system does anyone update the context. Six weeks later the team quietly stops trusting the drafts, and once trust goes, usage follows within a fortnight.

The fix: name the operator before you name the go-live date. Give them the trace view, the eval dashboard, and the authority to drop the autonomy level without asking permission.

3. It was rolled out as a tool, so it was treated as optional

If adoption is voluntary and the agent sits in a separate tab, you have built a thing people must remember to use. They won't, because the whole problem you were solving was that they're too busy.

The pilot team used it because they were invested. The wider team has a working process, a queue, and no spare attention for a second one.

An agent that lives inside the existing surface — the shared inbox, WhatsApp, the CRM record people already open — gets used because it's in the path. One in a separate portal gets used for three weeks.

The fix: integrate into the channel where the work already arrives, and make the agent's output the default starting point rather than an alternative to it. A draft that's already sitting in the reply box gets edited. A draft in another tab gets ignored.

4. The success metric moved when it got real

Pilots are measured on capability — accuracy, quality, "would you have sent this?" Rollouts get measured on business outcomes, and the two are connected by a chain of assumptions nobody wrote down.

So the agent drafts excellent replies, response time drops, and someone senior asks whether revenue moved. The honest answer is "we don't know, we never instrumented that," and in the vacuum the project gets recategorised as an experiment.

The fix: at pilot time, write down the one number that would justify the rollout, and instrument it before go-live so you have a baseline. Not a dashboard of twelve metrics — one number, agreed by the person who controls the budget.

The pattern underneath

All four are the same mistake wearing different clothes: treating the pilot as a technical question when it was always an operational one.

The technical question — can a model do this — was answered honestly in about a week, and the answer is usually yes. Everything after that is distribution, ownership, integration, and measurement. Which is unglamorous, and is exactly why so many organisations have a graveyard of successful pilots.

What a rollout-shaped pilot looks like

If you're about to start one, four things make the difference, and none of them are about the model:

  • Run it on a slice of the real distribution, ugly cases included, and count what you couldn't handle.
  • Name the operator and give them the controls on day one.
  • Ship it inside the existing surface, not beside it.
  • Agree one number and baseline it before you start.

Do those and the rollout is boring, which is the goal. Skip them and you get a great demo and a quiet six months.


Wednesday: RAG that survives a real corpus — what happens to retrieval when the documents stop being a tidy sample.

Got a workflow worth automating?

Tell us about it in a 30-minute call. If there's a fit, you'll get a scoped agent plan within 48 hours.

Book a call →