Research · 04 of 10 · Series · 02

Why Five Thousand Small Agents Beat One Large One

Failure analysis, not a taste for scale: why thousands of narrow agents beat one large one.

The most common question I get is why we run five thousand to fifty thousand agents per lifecycle instead of one capable agent with a long context window and a good set of tools. The answer comes from failure analysis, not from a preference for scale.

A monolithic agent has one property that makes it unacceptable in regulated work: it sees the whole job. That sounds like an advantage. In practice it means the agent's belief about the world and its permission to act on the world are held in the same place. When the belief is wrong, the action follows. There is no seam where a policy can intervene, no boundary where a failure stops instead of propagating.

Our design principle is the opposite. No single agent sees the whole trade. Each agent is scoped to exactly one gate in the lifecycle and does exactly one job. A signal agent reads NAV and flow data. A prediction agent estimates the direction and duration of a gap. A decision agent proposes one sized, routed action. None of them can execute anything. Their output is a proposal, and a proposal is just a data structure until something downstream permits it.

The verification argument

Small agents are verifiable in a way large agents are not. An agent whose job is to estimate borrow cost for a single basket can be tested against historical borrow data, given a bounded tool set, and assigned a permission scope that covers only what it needs to read. When it is wrong, we know which agent was wrong and at which gate. When a monolithic agent is wrong, we know only that the output was wrong.

This is the same reason distributed systems engineers moved from monoliths to services, and the same reason safety-critical software is built from small components with explicit interfaces. The agent layer is not exempt from those lessons because the components happen to reason in natural language.

The economic argument

The swarm size is also driven by the economics of the work itself. At five percent a year, a dollar earns roughly two ten-thousandths of a cent per day. A human desk cannot cover its own cost on a band that thin. Wall Street leaves the four to eight percent yield band in ETF primary markets alone for that reason, not because the opportunity is invisible.

A swarm changes the arithmetic. Thousands of narrow agents, each holding no inventory and each waiting on no counterparty, can work a band that is too small for people. The scale is not a demonstration. It is the only configuration in which the lifecycle is profitable to run.

What the swarm shares

The agents are small, but they are not isolated. Every agent reasons on the same model layer, a language model post-trained on the lifecycle's rules, documents, and exceptions. That shared substrate is what keeps five thousand independent decisions consistent with one rulebook. I will cover the post-training work in a later post. The point here is that specialization at the agent level and consistency at the model level are complementary, not in tension.

OpenEXA Research · Founder's notes · 04 / 10