The math of a parallel runtime
Why the choice of runtime foundation stops being a preference and becomes an infrastructure cost once volume arrives. And why it doesn't make runs any faster.
Tessera Engineering Team · Sep 14, 2026 · 7 min
The right question isn't speed
When people talk about agent performance, the first question is usually how long a run takes. It's the wrong question, and it's worth explaining why before anything else.
Most of an agent run's wall-clock time is waiting. Waiting for the model's response, waiting for the system being queried, waiting for the tool being called. No runtime speeds that up, because no runtime controls what's on the other side.
The right question is how many concurrent runs the same machine can sustain. That one is decided by the runtime, and it's the one that shows up on the invoice.
What changes with volume
An agent in production is a process that spends most of its life waiting and, in between, does short bursts of work.
With ten concurrent runs, any architecture works. With ten thousand, the question becomes what it costs to keep ten thousand things waiting at once.
If each run ties up an expensive unit of execution while it waits, the bill grows with the number of runs, not with the work done. You pay to wait. That's the point where the choice of foundation stops being a team preference and becomes a line item.
The Actor Model, and what it solves
The Tessera Runtime is written in Java 25, using the Actor Model. Two choices, and the second matters more than the first.
In the Actor Model, each run is an isolated unit with its own state that communicates by message rather than shared memory. That has three practical consequences.
Isolation by construction. A run that hangs doesn't hang its neighbors, because they share no state. In a system with thousands of agents from different teams, that stops being elegance and becomes a requirement: one department's badly written agent can't take down another's operations.
Parallelism without shared locks. The classic concurrency bottleneck is contention for a shared resource. Without shared memory there's no contention, and scale no longer depends on how well someone tuned a pool.
Density. An actor is a cheap unit of execution, and thousands of them waiting cost little. That's where the difference in how much hardware the same workload needs comes from.
Why Java, and not the obvious choice
Most of the agent ecosystem is Python, and for building and experimenting that's a real advantage: the libraries are there and the iteration loop is short.
In production at volume the math changes, for one specific reason: the platform sits on the critical path of every run. It isn't a service that answers a request and steps aside. It's the layer every call from every agent in the organization passes through, all the time.
In that position, sustaining high concurrency with predictable resource usage is no longer a detail. It's also the foundation the critical systems most of these companies have run for years are already built on, which means operations teams know how to monitor, tune and troubleshoot what's underneath.
About the number
This site publishes an internal measurement of a complete agent run in under 46 milliseconds, with the caveat that performance results depend on workload, hardware, configuration and methodology.
The caveat isn't a formality. It's the difference between publishing an engineering signal and publishing a marketing claim, and Tessera prefers the former. Any performance number that appears without its measurement conditions next to it should be read with suspicion, including ours.
What the runtime doesn't solve
It doesn't make the model faster. If a model call takes two seconds, it will take two seconds. The runtime's gain is in density and concurrency, not in the latency of a single run.
It doesn't fix a poorly designed agent. An agent that runs eleven searches where one would do will just run eleven searches faster. Bad design isn't an infrastructure problem, which is why observability matters as much as the runtime.
And it isn't the reason to buy. It's what holds the reason up: governance is only viable if the layer enforcing it doesn't become a bottleneck. Performance here is a precondition, not the pitch. Explore the Runtime.
