Research · 01 of 10 · Research note

Compounding Error and the Case for Decomposition. A Research Note on Agent Architecture

This post steps back from any particular product. It lays out the technical reasons, drawn from machine learning and systems research, why long-horizon autonomous work should be decomposed into many narrow agents behind a deterministic execution boundary rather than handed to one capable model.

The problem is horizon, not capability

Most discussion of agent reliability focuses on model capability. Bigger models, better reasoning, longer context. That framing misses the dominant failure mode in autonomous work, which is horizon length.

Consider an agent that must take a sequence of dependent steps to complete a job. Suppose each step is correct with probability p, and errors are not recoverable. The probability that a run of n steps completes correctly is p to the power n. At p equal to 0.99 and a hundred steps, the success rate is about 37 percent. At a thousand steps, it is effectively zero. No plausible improvement in per-step accuracy changes the shape of that curve. A lifecycle that runs from market open to settlement is thousands of steps.

This is not a new observation. In imitation learning, it is the compounding error problem. A policy trained to mimic expert behavior on the expert's state distribution drifts into states the expert never visited, where its errors are larger, which drives it further off distribution. Ross and Bagnell showed in 2010 that the expected cost of a naive behavior-cloning policy grows quadratically with horizon, and that interactive data collection (their DAgger algorithm) is needed to make it linear. Language models used as agents are behavior-cloning policies at scale. They inherit the quadratic term.

The autoregressive structure of the models makes this worse. Every token is conditioned on the model's own previous outputs. A hallucinated intermediate fact becomes context for every subsequent step. The literature calls this exposure bias, and it means a monolithic agent has no natural point at which an error stops propagating.

Decomposition changes the exponent

The standard engineering response to a p-to-the-n problem is to shorten n and insert checkpoints. If a job of a thousand steps is decomposed into a thousand single-step agents, each of which produces an output that is verified before the next agent consumes it, the failure model changes. An error at step k is caught at the verification gate after step k. It does not become context for step k plus one. The success probability of the run is no longer the product of a thousand unverified steps. It is the product of a thousand steps each with a verifier, and a verifier only has to be good at one narrow question.

This is why the number of agents in a well-designed system is large. It is not parallelism for throughput. It is decomposition for error containment. Each agent's scope is chosen so that its output can be checked by a simple rule or a narrow model, and so that its failure is local.

There is a second benefit. A narrow agent has a narrow input distribution. The state space a borrow-cost estimator sees is small and stable. The state space a do-everything agent sees is the whole world. Narrow distributions are the regime in which learned models are reliable and in which their errors are measurable. Distribution shift, the thing that kills deployed models, is bounded by construction when the agent's job is bounded.

The generator-verifier gap

A further reason to separate proposing from approving comes from a well-documented asymmetry. For many problems, verifying a candidate answer is far easier than producing one. This is the intuition behind the P versus NP distinction and it shows up empirically in language models: a model asked to check a proposed solution against explicit criteria is markedly more accurate than the same model asked to produce the solution unaided. Recent work on process reward models and verifier-guided search exploits exactly this gap, training a separate verifier to score intermediate steps rather than trusting the generator's own confidence.

An execution council in an agent architecture is a verifier made explicit. It does not need to know how to trade. It needs to know the policy, the caps, the hours, and the conflicts, and it needs to check a structured proposal against them. That check is a small, well-posed problem. The generator can be probabilistic and occasionally wrong. The verifier is where reliability is concentrated.

The corollary is that the verifier should not be the same model instance as the generator. Self-verification by a single model inherits the model's own blind spots. Independent verification, whether by a rules engine or a separately trained model, does not.

Why the execution layer should not be a language model

Language models are poor at a specific class of operations: those that require exactly-once semantics, exact arithmetic on identifiers, and deterministic state transitions. The reasons are structural. Sampling introduces variance even at low temperature. Tokenization fragments long numeric strings and identifiers in ways that make exact reproduction unreliable. The model has no persistent state between calls except what is placed in context, so it cannot natively guarantee that an action already taken is not taken again.

Distributed systems research solved exactly-once execution decades ago, through idempotency keys, write-ahead logs, and reconciliation against an authoritative source. Those solutions are deterministic code. The correct architecture uses the language model for what it is good at, interpreting ambiguous state and proposing structured actions, and hands the structured action to a deterministic state machine for execution. The boundary between them is a typed schema. The model emits a proposal that conforms to the schema or it emits nothing that executes.

This is the neuro-symbolic division of labor, and it is older than the current wave of agents. Perception and reasoning are learned. Action is compiled. The interesting research is in making the interface between them tight enough that nothing leaks across.

Capability-based permissions as the enforcement mechanism

Decomposition and verification contain errors. They do not by themselves prevent a wrong-but-approved action from doing damage beyond its intended scope. That requires a permission model.

The right model comes from object-capability security. An agent does not have an identity that is checked against an access list at execution time. It holds capabilities, unforgeable tokens that grant specific rights to specific resources, and it can only invoke what it holds. A prediction agent holds a read capability on market data. It does not hold, and cannot acquire, a capability that places orders. This is the principle of least privilege applied at the granularity of individual agents and individual tools, and it is enforced by the runtime rather than by the model's good behavior.

The practical consequence is that the blast radius of any single agent is fixed at design time. The model can be jailbroken, confused, or wrong. It still cannot call a tool it does not hold a capability for. Combined with a deterministic execution layer and an independent verifier, this is what allows a system of thousands of probabilistic components to be safe in aggregate.

Summary

The argument for swarms of narrow agents is not a preference for scale. It follows from four results. Compounding error grows with horizon, so horizons must be shortened with verification gates. Verification is easier than generation, so verifiers should be separate and explicit. Language models cannot guarantee exactly-once execution, so execution must be deterministic code. And permission must be enforced by the runtime, not by the model, so capabilities must be scoped per agent. An architecture that respects all four looks like many small agents, one council, one state machine, and one permission layer. That is not a product decision. It is what the research implies.

OpenEXA Research · Founder's notes · 01 / 10