AI agent governance: a guide for banks

Governing agents is not writing a policy or building an approval queue. It is being able to answer five questions about any agent, at any time, without opening an investigation.

Tessera Engineering Team · Oct 6, 2026 · 7 min

The risk committee asks

The risk committee asks for something simple: a list of the AI agents in production, who owns each one, and what each one can reach.

The answer takes two weeks. Customer service runs three agents. Lending has one nobody calls an agent, because it started life as a script. A collections pilot has been talking to real customers since last month. And there is a fifth, built by a vendor, running in an environment nobody in IT administers.

None of these agents is doing anything wrong. The problem is that nobody can say so without assembling a task force. That is the symptom AI agent governance exists to fix.

What governance is not

The word attracts three answers that look like governance and are not.

A document. The board-approved AI use policy says what is allowed. It does not stop an agent from calling a system it should not, because the agent never reads the document.

An approval queue. If every prompt change needs three signatures, teams stop changing prompts through the official path. The queue becomes the reason to route around the rule.

A dashboard. Seeing that an agent made two hundred thousand calls yesterday is observation. Governing means having decided beforehand what it was allowed to call.

For anyone who has to account for it, governance is the ability to answer questions about any agent without relying on the memory of whoever built it.

Five questions

Which agents exist. Not the ones the committee approved, but the ones actually running: which environment, which version, since when. An inventory kept by hand goes stale the day someone ships without telling anyone.

Who owns each one. A name, not a team. When the agent does something unexpected at eleven at night, someone has to be paged, and that someone needs the authority to switch it off.

What each one is allowed to do. Which systems it can call, which data it can read, which models it can use, how much it can spend, and what needs a person before it proceeds. The list has to be the rule in force, not the intent of whoever designed the agent.

What changed, and who changed it. Prompt, policy, configuration and version, with a date and an author. Most agent incidents start with a small change that looked harmless.

What happened in one specific run. When a customer complains about an answer, the question is about a case, not an average. The answer has to be a record of that case.

If all five can be answered in minutes, the operation is governed, even if the policy is still short. If any of them needs a meeting, that is where the gap is.

Where the rule lives

The single most important architecture decision is this one: does the rule live inside each agent, or outside all of them.

Inside each agent feels faster. Every team writes its own sensitive data filter, its own spend cap, its own permission check. It works for the first agent. By the tenth there are ten versions of the same rule, in different frameworks, and the auditor has to check all ten.

Outside all of them, the rule is evaluated in a shared layer before the call leaves: if it blocks, the call never goes out; if it requires approval, the run waits for a person; if it allows, the call proceeds and the record is kept. The agent still chooses the path. The boundary still belongs to the institution.

In a bank, where agents arrive from internal teams, vendors and several platforms, only the second option scales. Not because it is more elegant, but because the examiner does not ask about a team. They ask about the institution.

What is different in a bank

The five questions apply to any company. In a bank, three things make improvising the answer more expensive.

Evidence is requested from outside. Internal audit, external audit and supervisors are not satisfied with "we believe the agent did not do that". They ask for the record.

Model risk practice already points this way. Banks used to supervisory guidance on model risk management, such as SR 11-7 in the US, already keep a model inventory with owners and validation. Agents raise the same questions with more moving parts: prompts, tools, and several models per run.

Agents do not come from one place. Some are built in house, some come from vendors, some existed before anyone called them agents. Governance that only covers what was built in the official tool misses exactly what is hardest to see.

Which rules apply to each use case is a question for each institution’s legal and compliance teams. What governance delivers is the ability to show, with a record, that the rule they defined was applied.

Where to start

Do not start with the full policy. Start with the questions that have no answer today.

First, the inventory. Register every agent that talks to a customer, a system or company data, including vendor agents and the ones that began as scripts. Registering is not migrating: the agent stays where it runs.

Then, one owner per agent. A name, with the responsibility written into the record, so it does not disappear when the person changes roles.

Next, a single rule, applied to everyone. Pick the one your CISO worries about most, usually the one that stops an agent from exposing personal data, and enforce it across the whole fleet before writing the second. One rule applied everywhere is worth more than twenty applied wherever a team remembered.

Finally, a record of every run, complete enough to answer a question about one case without rebuilding it from logs.

What governance does not solve

It does not make the model right. A well governed agent can still give a bad answer; the difference is that the answer is recorded, within bounds, with an owner who can fix it.

It does not replace human judgment about what to automate. Which journeys an agent may take on is a business and risk decision. Governance makes sure the decision is enforced, not that it is a good one.

And it has a cost. Every call evaluated against a rule goes through one more step. That cost is small next to answering the risk committee two weeks late, but it is real, and worth measuring.

Central governance, with teams shipping at their own pace inside limits the institution defined. That is the design Tessera implements. Explore Governance.

Govern what you've already built.

Connect agents from different frameworks to a common layer for operations and governance.

By submitting, you agree to receive the Tessera newsletter, as described in the Privacy Policy. You can unsubscribe at any time.