Research · 02 of 10 · Research note

Specializing a Model to a Rulebook and Proving What It Did. A Research Note on Post-Training and Audit

This post covers two technical questions that any autonomous system in regulated work must answer. How do you make a general language model reliably follow a specific rulebook? And how do you prove, after the fact, exactly what the system did and why?

Part one: retrieval is not enough

The default approach to giving a language model domain knowledge is retrieval-augmented generation. Store the documents, retrieve the relevant passages at inference time, place them in context. It works well for question answering. It works poorly for rule-following under exceptions, for three reasons.

First, retrieval selects by similarity, and the rule that governs an exception is often not textually similar to the exception. A shortfall in a basket delivery after cutoff is governed by a procedures clause that mentions neither shortfall nor cutoff in the same sentence. Embedding similarity misses it.

Second, retrieved context competes with the model's prior. A general model has strong priors about how finance works from pretraining. When a retrieved rule contradicts the prior, the model does not reliably defer to the rule. Studies of knowledge conflicts between context and parametric memory consistently show the model splitting the difference or ignoring the context under adversarial phrasing.

Third, retrieval does not teach the model the mapping from messy state to rule application. The documents say what the rule is. They do not say what the world looks like when the rule applies. That mapping is exactly what a human specialist has and a general model lacks.

Post-training moves the rulebook into the weights

Post-training, meaning supervised fine-tuning followed by preference optimization on domain data, changes the model's prior rather than competing with it. The rulebook becomes what the model believes, not what it is told.

The supervised stage is straightforward in principle. Each training example is a lifecycle state, the applicable rule, and the correct structured action. The hard part is the data. Rules are documented. Correct actions under exceptions are not. They live in resolution logs, reconciliation records, and the memory of the people who handled them. Building the training set means turning every historical exception into a case with its state, its resolution, and its confirmation. The append-only ledger discussed later in this post is, among other things, the mechanism that generates this data continuously once the system is live.

The preference stage matters more than it appears. Direct Preference Optimization and its relatives train the model to prefer one response over another given the same state. In a governed system, the natural preference signal is the verifier. Every proposal the council rejected is a negative example paired with the state that produced it. Every proposal the council approved and the counterparty confirmed is a positive example. The model learns where the policy boundary is without anyone writing it down as a rule, because the verifier's decisions are the labels.

This creates a feedback loop that is unusual in deployed machine learning. Most systems train once and drift. A system with an explicit verifier and an append-only record of its decisions generates labeled data every session, and the labels come from the component that is by construction the most reliable in the system.

Parameter-efficient specialization and the forgetting problem

Full fine-tuning of a large model on a narrow domain risks catastrophic forgetting: the model gets better at the rulebook and worse at general reasoning it still needs. Parameter-efficient methods such as low-rank adaptation train a small set of additional weights while freezing the base model. The base model's reasoning is preserved. The adapter carries the domain.

This has an architectural consequence for multi-lifecycle systems. One base model, many adapters. A new lifecycle is a new adapter trained on that lifecycle's rules and exceptions, sharing the base with every other lifecycle. The replication of a proven agent onto a new asset class becomes, at the model layer, the training of one adapter on one new case set while the reasoning substrate stays fixed. That is the technical basis for a master-and-copy agent model.

Constrained decoding closes the schema gap

A post-trained model that reasons correctly must still emit an action the execution layer can consume. Free-form text is not an action. The solution is constrained decoding. The model's output is restricted at the token level to conform to a grammar or schema, typically a JSON schema describing the proposal type, its fields, and their permitted values. Tokens that would violate the schema are masked before sampling.

This is not a formatting convenience. It is the mechanism by which the boundary between probabilistic reasoning and deterministic execution is made airtight. The model cannot emit a proposal with a field the schema does not define, a venue the schema does not permit, or a size outside the schema's range. Whatever it believes, its output is a well-typed structure or it is nothing. The execution layer never parses natural language.

Part two: proving what happened

A system that acts autonomously in regulated work must be able to answer, for any past moment, what it did, why, and on what authority. That is an audit requirement and it is also a research requirement, because it is the only way to evaluate and improve the system offline.

The data structure for this is a hash chain, the same primitive that underlies Merkle trees, version control systems, and certificate transparency logs. Each record includes a cryptographic hash of its own contents and the hash of the previous record. The chain has one property that matters: any modification to any record changes its hash, which invalidates the next record's stored reference, which cascades to the end of the chain. Tampering is not prevented. It is made detectable by anyone holding the final hash.

Append-only is the complementary property. The store accepts new records and refuses updates and deletes at the storage layer, not merely by policy. Together, the two properties mean the record is a fact about the past that no administrator can revise.

What must go in the record

The content of each record is as important as its integrity. For an autonomous system, a useful record captures the agent identity and permission scope that produced a proposal, the full structured proposal, the verifier's decision and the specific policy check that determined it, the execution result, and the independent counterparty confirmation. Recording the counterparty confirmation as the condition for the settled record, rather than the system's own belief that an action completed, is what makes the ledger ground truth rather than self-report.

Per-agent attribution is what makes a swarm accountable. When thousands of components act concurrently, the question after an incident is never what the system did. It is which agent, at which gate, under which permission, produced the proposal that led to the outcome. A ledger keyed by agent identity answers that directly.

Replay as off-policy evaluation

The research payoff of a replayable ledger is that every live session becomes an evaluation environment. Because the record contains the full state at each decision point and the action taken, one can ask what a different policy would have done against the same states. This is off-policy evaluation, a well-developed area in reinforcement learning, and the ledger is what makes it possible without simulation.

Concretely: a new model checkpoint can be run against the recorded states of every past session and its proposals compared to the verifier's historical decisions before it touches live capital. A tighter council policy can be evaluated by counting which historical proposals it would have rejected and what those proposals subsequently did. A change to an agent's permission scope can be checked for whether it would have blocked any action that was in fact confirmed and profitable. None of this requires a market simulator, which is fortunate, because market simulators are the least trustworthy component in any quantitative pipeline.

Summary

The two halves of this post are one argument. Post-training moves a rulebook into a model's weights, with the verifier's decisions as the preference signal and constrained decoding as the guarantee that outputs are well-typed. The append-only, hash-chained ledger records every decision with per-agent attribution and counterparty confirmation, which serves audit and simultaneously generates the training data and the evaluation environment for the next model. The model gets more specific to the lifecycle the longer it runs. The record of what it did is the reason it can.

---

Disclosures: Performance figures referenced in this series are proof-of-concept results over a limited period and are not guarantees of future results. OpenEXA does not take custody of client assets. Lifecycles beyond Lifecycle 01 describe roadmap intent. Any offering of securities would be made only to accredited investors under Rule 506(c). Named brokers and exchanges do not endorse OpenEXA.

OpenEXA Research · Founder's notes · 02 / 10