Where policy is evaluated, and why it happens first
The difference between observing and governing is when the rule runs. Before the call goes out is a different thing from after, and the cost of that choice is real.
Tessera Engineering Team · Sep 14, 2026 · 6 min
The objection, which is fair
Every time someone proposes a governance layer for agents, the first reaction from the people who'll run it is the same: isn't that just one more thing in the way?
It's a fair objection, and it deserves the technical answer, not the sales one.
The short answer is that the layer already exists. It's spread across spreadsheets, policy documents nobody reads, the tacit knowledge of whoever built the agent, and the manual reviews of whoever approves it. What changes isn't whether it exists. It's where it runs.
The long answer is this article.
Two possible places
An organizational rule about what an agent can do can only be applied at two moments: before the action happens, or after.
After is what most tools offer, and they call it observability. The agent calls the tool, accesses the system, spends the token, and the record shows up. If the call shouldn't have gone out, you find out that it did.
That has value, and it isn't governance. It's bookkeeping.
Before means the call is evaluated before it exists. The agent decides it wants to call a tool, that intent is checked against the organization's rule, and only then does it become an actual call, or not.
The practical difference is the same one that separates a firewall from a traffic report. Both know what got through. Only one decides what gets through.
The sequence, step by step
In Tessera, this is the path a call takes.
Midway through a run, the agent decides it needs something from outside: to call a tool, fetch data, use a model. That decision is the agent's, and it's the part that makes the agent useful.
The intent is intercepted before it becomes a call. At that point we have the agent, its version, the environment, the declared owner, and exactly what it's trying to do.
The policies that apply to that scope are evaluated. Not every policy in the organization: only those whose scope covers it, which can be by agent, by team, by environment or by resource type.
Then there are three outcomes. The rule allows it, and the call proceeds, with a record of what was evaluated. The rule blocks it, and the call doesn't go out, with a record of which rule blocked it. Or the rule requires human approval, and the run pauses and waits, which is the most interesting of the three.
The pause is the hard case
Blocking is simple to implement and simple to explain. Pausing and waiting isn't.
A run that pauses because policy requires approval needs state: someone has to be notified, the run has to survive the wait, and when the approval comes in it has to resume from exactly where it stopped, not start over.
That's why this capability lives in the platform layer and not inside the agent. An agent that implements its own pause is implementing execution-state persistence, and every team that does it will do it differently.
What this costs
Here's the part the opening objection was right to raise.
Evaluating policy before the call happens on the critical path. Every tool call and every model call goes through it. That costs time, and there's no point pretending it doesn't.
What we can say honestly is what determines that cost. Evaluation is local, doesn't go over the network, and works on a rule set that's already loaded. It doesn't depend on the model, doesn't depend on the target system, and doesn't grow with the size of the response.
And there's a comparison the objection forgets. The alternative to evaluating before isn't not evaluating: it's evaluating after, by hand, when someone asks. That cost doesn't show up in the run's latency. It shows up in the hours of whoever has to investigate.
What policy evaluation doesn't solve
It doesn't decide what the right policy is. A policy engine enforces what the organization defined, exactly as precisely as it was defined. If the rule is wrong, it will be enforced correctly and the result will be wrong at scale, which is worse than having no rule at all.
It doesn't look inside the model's response. Evaluation happens on the intent to call: which tool, which arguments, which model, which data. It doesn't judge whether the text the model returned is appropriate. That's a different problem, with different tools, and confusing the two is the most common mistake in conversations about agent security.
And it doesn't replace the record. Policy decides what happens; the trail keeps what happened. They're two different things, and whoever has only one will discover the missing one at the worst possible moment.
Governance enforces the rule before the call goes out. Audit & Replay keeps what happened afterward. Explore Governance.
