Why AI agent pilots stall before production

The pilot proved the agent works. What blocks the move to production is almost never the model. It is the questions the pilot never had to answer.

Tessera Engineering Team · Oct 20, 2026 · 7 min

The demo went well

The demo went well. The agent read the request, checked two systems, got the answer right, and leadership left the room wanting it live next quarter.

Three months later it is still a pilot. Nobody gave up on it. It is stuck in a string of meetings: security wants to know what it can reach, audit wants to know how a case gets reconstructed, finance asks what it will cost at real volume, and operations wants to know who picks up when it fails at 3 a.m.

None of those questions is about whether the agent works. All of them are about what happens when it stops being a demo and becomes operations. That is where most AI agents stall on the way to production.

What the pilot never had to answer

A pilot is designed to prove value. It runs on hand-picked data, with users who know it is a test, and a team watching every result. That is the right call: it is the cheapest way to learn whether the idea holds up.

The side effect is that everything not needed to prove value gets left out. A hard-coded credential, because it was only a test. Logs in whatever format the developer found useful. No spend limit, because volume was small. No formal owner, because the owner was the whole team, and the team was watching.

When it is time to switch it on for real, each shortcut turns into a question. And since nobody planned to answer them, they get answered one at a time, in meetings.

Four questions

What can it reach. Which systems, which data, under which identity. In the pilot, the agent ran on its builder’s credentials. In production, security needs a dedicated identity, least privilege, and assurance that it will not read what it should not, including when someone tries to talk it into it.

How do you explain a case. When a customer disputes an answer, someone will have to show which agent version ran, with which prompt, which model, which rules were evaluated, and what the tools returned. A debug log does not answer that.

What does it really cost. Pilot cost says nothing about production cost. A run that calls the model four times instead of once, multiplied by real volume, changes the math. Finance wants cost per agent and per case before approving, not on next month’s invoice.

Who runs it. Who gets paged when it fails, how a new version ships without taking the old one down, how you switch off an agent that starts behaving oddly, and how fast.

Why the second agent is not faster

Say the team answers all four. It sets up the identity, builds the trail, adds a spend cap, defines the on-call rotation. The agent goes live.

Then the second agent arrives, from another team, built on another framework. And the four questions come back from zero. The first team solved them inside its own project, in whatever way fit that codebase. None of it carries over.

That is why so many companies have one or two agents in production and a queue of stalled pilots behind them. Getting each one through the door costs almost as much as the first, and the door always has the same meetings.

What changes when the answers live in one place

The four questions do not change from one agent to the next. The agent does. That suggests the answers should not live inside each project, but in a shared layer underneath all of them.

In that layer, a new agent is born registered, with an owner and a version. Access and data rules already apply to it, because they belong to the organization, not the team. Every run is recorded in the same format as every other. Cost shows up by agent, team and model from the first call.

Teams still choose the framework, the model and how to build. What they stop doing is reinventing the part that makes their agent no different from anyone else’s.

Pilots prove value. Production requires infrastructure. The line is simple, and the practical consequence is concrete: the question for the next pilot stops being "how do we get this into production" and becomes "what is missing for it to run on the layer we already have".

A test before the next pilot

Before approving the next pilot, ask the four questions in reverse. If this agent works, do we know which identity it will use in production? Do we know where the record of each run will live? Do we know how much it can spend a month, and who sees that? Do we know who gets paged when it fails?

If the answers are "we will figure it out later", the pilot will probably work, and will probably stall exactly where the last ones did.

What infrastructure does not fix

It will not rescue a weak use case. If the agent does not solve a problem someone pays to have solved, production only makes that visible sooner.

It does not replace the risk decision. Which journeys an agent may take on is still a call for the business and for whoever owns the risk. The layer makes sure that call is enforced, recorded and measured.

Tessera AgentOS brings governance, execution, observability and development into one layer, so the next agent does not have to sit through the same meetings as the first. Explore Tessera AgentOS.

Govern what you've already built.

Connect agents from different frameworks to a common layer for operations and governance.

By submitting, you agree to receive the Tessera newsletter, as described in the Privacy Policy. You can unsubscribe at any time.