What goes into an audit trail

When someone asks what an agent did, the answer has to be a record, not a reconstruction. This is what gets kept, and what doesn't.

Tessera Engineering Team · Sep 14, 2026 · 6 min

Two weeks later

An agent granted a discount the policy didn't allow. Or approved an exception, or shared information the company wouldn't have. Nobody noticed at the time, because at that moment it was just one of the day's thousands of runs.

Two weeks later it turns into a complaint, and someone has to say whether it happened.

This is where most operations find out what they have and what they don't. The answer starts in a logging tool: you search by timestamp, cross-reference the session ID, and try to piece the sequence together from lines written to diagnose infrastructure, not to explain a decision.

When the answer arrives, it's a reconstruction. And a reconstruction isn't evidence.

The difference between a log and a trail

An application log records the events the developer thought were worth recording, in whatever format made sense in that file, at whatever level of detail seemed sufficient. It's great for finding out why a service went down at three in the morning.

An audit trail records what's needed to explain a decision, completely and at the moment it happened. The difference isn't volume, it's intent: the log answers "is the system healthy," the trail answers "why was this decided this way."

The consequence shows up in the question the log can't answer: which version of the prompt was live in that run. The log records that a model call happened. It doesn't record which configuration produced that text, because whoever wrote the log line wasn't thinking about accountability.

The eight fields

For every run, the trail keeps eight things. It isn't a feature list: it's the minimum set needed to answer a question about a case.

The agent and its owner: which agent ran, in which environment, and who owned it at that moment. The owner is recorded with the run rather than read from the current registry, because the person may have changed teams since then.

The exact version: the revision that ran, not the one that's live today. If the agent was changed afterward, the record still points to the one that produced that result. This is the field most often missing from homegrown implementations, and the most expensive one to lack.

The prompt as it was sent, with its variables already resolved, not the template that generated it. A template with three conditionals produces different texts, and it's the text that has to be kept.

The model and its parameters: which model responded, at what temperature, with what context limit and through which provider, for every call. Models get deprecated and replaced; the trail has to say which one was there.

The policies evaluated: every rule considered, the result of each, and which one blocked when something was blocked. This is the field that turns the trail from descriptive into explanatory. Without it you can know what happened, but not why it was allowed.

The tools called: which calls went out, with which arguments, and what came back. Including the ones that failed, because silent failure is often the cause of strange behavior.

Consumption: tokens and cost per step, attributed to the agent, the team and the environment. It's in the trail and not only in the cost report because the question "why did this run cost ten times more" is about a single case, and single cases are where the trail lives.

The outcome: what happened in the end, and whether it matches what was requested.

Replay, don't reconstruct

With all eight fields, you can do something log-based reconstruction never allows: run it again.

The replay starts from the checkpoint where the decision was made, with the same configuration, the same prompt and the same policies. It isn't a rough simulation of what probably happened. It's the run, again.

That changes the nature of the conversation with whoever is asking. Instead of "we believe the agent considered X," the answer becomes "look: with this rule and this prompt, this is the result."

And it changes what you can do afterward. Once the policy is changed, the replay shows whether the outcome would have been different. The trail stops being just a defense and becomes a tuning instrument.

What the trail doesn't solve

Three limits, stated here because you'd discover them anyway.

It doesn't explain the model. The trail says which model responded, with which parameters, and what it returned. It doesn't say why that model produced that output, because that lives inside the model, not the platform. Anyone promising LLM decision explainability from an execution trail is promising something else.

It doesn't replace policy. A complete trail of a poorly governed operation produces an excellent record of bad decisions. The record is what lets you find out; it isn't what prevents. That's why policy evaluation happens before the call, and the trail after.

It has a storage cost. Keeping the resolved prompt and tool responses for every run of a large fleet isn't free, and the retention policy is the organization's decision, with different retention periods depending on how critical each agent is.

Why this becomes infrastructure

The reasonable question is why this doesn't live inside each agent, handled by whoever built it.

The answer is the same as for every platform decision: because then there would be as many trail formats as there are teams, and "what happened in that run" would have a different answer in every department. When the auditor asks, they don't ask about a team. They ask about the organization.

An audit trail is the kind of thing that only works when it's the same everywhere. And whatever needs to be the same everywhere belongs in the layer underneath.

Governance enforces the rule before the call goes out. Audit & Replay keeps what happened afterward. Explore Governance.

Govern what you've already built.

Connect agents from different frameworks to a common layer for operations and governance.