AI Eval against Causal Generators

A right answer can describe impossible dynamics. Causal generators test AI claims against validated models of mechanisms, with explicit limits and room to say inconclusive.

By E. M. AbdullahPublished Updated

Paper and code: the working paper and its Python release are on Zenodo at doi:10.5281/zenodo.23060990.

Language models are increasingly asked to forecast, plan and explain systems that change over time: inventory, queues, populations, capacity. Most of the evals we run on that reasoning check the final answer, or ask another model whether the answer looks right. For systems like these, that misses something important. An answer can be right while the path it describes is one the system could never follow.

Here's a simple case, picture a simple fishery model, with max capacity of 100 fish and constant harvesting. Above harvest 7.5, its population eventually collapses. Set the harvest to 8 and ask three forecasts whether the population collapses. All three might correctly say yes, but they may describe different paths. One follows what the model actually does: a decline that slows near half capacity before accelerating into collapse. One is a steady straight line down. One tracks the true path with a small drift.

An answer check passes all three. But anything you build on the straight-line forecast inherits dynamics that don't actually follow how the fishery works. In the paper we call that kind of claim dynamically inconsistent: the answer is right, but the trajectory doesn't follow the real mechanism.

At harvest 8, compare the straight line with the model's trajectory, then sketch a forecast of your own. Which checks change?

One fishery, three claims
Harvest set to 8, just above the sustainable limit of 7.5. Only the true trajectory follows the system's flow.

What you're scoring against is a causal generator: a small, executable model of how a system changes over time and how it responds when something is set or changed. It has three parts: the state (the fish population), the interventions you can set (the harvest), and the flow, the rule for how the state changes given both (growth toward capacity, minus the harvest). Give it a starting state and an intervention, and it produces the trajectory it predicts. A claim can then be checked against that prediction. It's a model recovered from data and validated over a declared range, so it's a fallible reference, not ground truth. The aim is a model you can read and inspect; as you'll see later, recovery doesn't always produce one.

If you've followed recent AI-for-math work, "generator" may already have an alternate meaning for you. Systems like DeepSeekMath-V2 pair a proof generator with a verifier, an LLM trained to check the proof, and they start from the same observation as this essay: a correct answer doesn't guarantee correct reasoning. A causal generator plays the other role. It's the reference a verifier checks against, an executable model of how a system actually behaves under intervention. The two ideas fit together. In that type of setup a causal generator can be the grounded verifier the system independently checks against.

Context is King

Most machine-learning models learn associations: what follows what in the data they were trained on. That's powerful, but causation needs more context. Pearl's distinction is between seeing an input take a value and setting it. The two generally differ, and data about the first doesn't automatically tell you about the second.

The fish-tank example makes this concrete. Two tanks each hold 10 litres. Tank A fills at 2 litres a minute and drains 20% of its volume each minute. Tank B fills at 4 litres a minute and drains 40%. Left alone, both hold exactly 10 litres forever. Their steady-state volume records are identical. Those records alone cannot identify their different responses to reduced inflow. Now cut each inflow by a litre a minute. Tank A settles at 5 litres. Tank B settles at 7.5.

The difference lives entirely in how each responds to change. Without additional information, an evaluator cannot choose between those responses reliably.

That makes this a data problem before it's a modelling one. Production systems are very good at logging outputs and metrics, and our analytics pipelines map the correlations well. What they rarely capture on purpose is the interventional record: what was changed, when, by how much, and under which conditions, with enough state around it to see the response unfold. Enterprises have done work at this scale before, moving to microservices, event streams and warehouses. Capturing interventions is a similar effort: treat them as first-class events, stored by intervention and regime, not only by time and entity.

Feedback is common in these systems. Orders respond to inventory, which responds to orders. The field has tools for this, including structural causal models with cycles and models derived from differential equations. A generator sits naturally alongside them, because it represents the dynamics directly as a flow over time, so feedback, delay and thresholds are part of the model rather than exceptions to it.

Building the evaluator

Build in this order: the data, the generator, claim capture, then scoring.

Get data that includes interventions. Intervention diversity matters more than volume. Record runs under different settings, with each intervention stored explicitly. Cover the range you'll test: in the fishery, four harvest levels weren't enough to separate the causal terms, and eleven were. Hold some regimes back entirely, including stressed ones, for validation. Write down the range the data covered, because that's the range the evaluator is allowed to speak about. And encode the hard limits: a stock can't go negative, so the generator and the scorer should both know that.

This connects to the Foundation patterns. Causal Lineage records the strategy chosen, the context compiled and the action taken, which is good raw material for intervention records. It still needs the relevant state, timing and variety of regimes; an audit trail alone doesn't identify a mechanism.

Recover the flow. Decide the causal frame first (which quantities are states, which are interventions, which are fixed) and then recover equations inside it. The prototype uses SINDy, the sparse-identification method from Brunton and colleagues, but it's one choice among several: state-space identification, symbolic regression and neural differential equations each make different assumptions. Whatever you use, normalise states by their natural scales first, or a small coefficient with a big effect (like crowding near capacity) can vanish from the fit. For the fishery, the recovered coefficients landed close to the true ones (0.3007 against 0.3, 1.0011 against 1.0), plus one small extra term, with a median rate error of 0.001 to 0.004 on harvests it had never seen.

Then validate it on the held-out regimes, against a threshold you set in advance. Passing tells you the model agrees with the system over the tested range. It doesn't prove you've found the true mechanism, and it says nothing about regimes you didn't test.

Capture claims in a structure you can score. Ask the model for a short answer plus its trajectory at stated times, in a fixed JSON schema, and audit a sample of the extractions. Otherwise you may end up measuring parsing errors instead of reasoning.

Score each claim four ways: does it break a hard rule, like a negative stock; does each step follow the model's flow from the claimed state; how far is it from the generator's prediction under the same intervention; and is the answer consistent with the tipping condition over the declared horizon?

Three Eval Outcomes

Every claim gets one of three outcomes: consistent, inconsistent, or inconclusive. Inconclusive should mean the result falls in a declared uncertainty bound or that the intervention is outside the range the generator was validated on. Tolerances are set before any scoring.

Inconclusive is a result, not a gap. Forcing a pass or fail where the evidence can't decide produces confident verdicts that cannot be trusted. The important note is that generators can be constructed to abstain correctly, this can be a powerful tool to help eval and prevent confident confabulations, so track how often your generator abstains, report it alongside pass/fail, and ensure it meets the standards to valid abstention.

In the fishery example, answer matching accepts four of five constructed claims. The generator accepts one, rejects two, and declines to decide on two: one because its drift sits inside the uncertainty band, and one because its harvest of 12 is outside the range the generator was validated on.

What the benchmark showed

To see where this holds up, we built a six-level complexity ladder: a single tank, the fishery, inventory with a shipping delay, an SEIR epidemic, a multi-delay supply chain (the beer game), and the chaotic Lorenz system. Each has an exact mechanism for labels. Across 612 claims and 6,108 judge calls, Sonnet 5, Haiku 4.5 and a local llama3 8B judged claims in three conditions: a description only, the exact equations, or the equations plus the generator's predicted trajectory. Every call ran with tools disabled. Lorenz labels also used empirical attractor screens, with forecast accuracy assessed over about one Lyapunov time.

Judging results balanced accuracy · 612 claims · 0.50 is chance

Judge and informationTankTank: one state, linear, one equilibrium approached smoothly. It’s the baseline: can the model reason about rates at all? A model that fails here can’t handle anything above it.FisheryFishery: still one state, but with nonlinear growth and a harvest threshold. Below the maximum sustainable harvest the population settles; above it, it collapses. It tests recognizing a tipping point and the right path shape. This is where right answers with impossible trajectories first appear.InventoryInventory with a shipping delay: two states (stock on hand, plus goods in the pipeline). Delay is the new ingredient, which brings overshoot and oscillation. It tests accounting for what’s ordered but not yet received, a well-documented human error that models can repeat.SEIRSEIR epidemic: four coupled nonlinear states. The epidemic depletes its own target population, so the peak’s timing and height aren’t obvious from any one equation. It tests coupled feedback. This is where the tested frontier models’ forecasts broke down.Beer gameBeer game: multiple stages combining level 3’s delays with level 4’s coupling. It produces the bullwhip effect, where small demand changes amplify upstream. It tests amplification through chains of delayed feedback, where each sensible local decision adds to the system-wide instability.LorenzLorenz: chaos (σ = 10, ρ = 28, β = 8/3). Point forecasts lose accuracy on the scale of a Lyapunov time, about 1.1 time units. It tests something no other level can: judging whether a trajectory is consistent with the dynamics, rather than whether it matches a single predicted path.
Checks without a model
Answer match0.600.640.580.560.620.68
Rules0.780.690.770.670.520.59
Reference lookup0.770.900.700.650.700.57
Sonnet 5
Description only0.940.740.860.940.860.90
With equations1.000.990.980.940.820.86
With generator output1.000.990.990.941.000.96
Haiku 4.5
Description only0.950.710.870.790.680.70
With equations1.000.860.870.680.660.75
With generator output1.000.920.910.830.900.93
gpt-oss:20b (local, open weight)
With generator output*Run after the main benchmark, on a stratified 120-claim subset (20 claims per level, half valid), to test how a local open-weight model with stronger reasoning compares when given the same instrument. Generator-output condition only, so it shows assisted performance, not how much the generator improved it. Run through Ollama on a consumer desktop with a 12 GB GPU, and not part of the pre-registered judge set. Timeouts were longer (600 and 900 s) because it reasons at length; 16 calls were retried once for output length and one unparsed verdict was counted as wrong. On those same 120 claims it scored 0.89 overall, against 0.93 for Haiku 4.5 and 0.97 for Sonnet 5, and caught 97% of invalid claims. Its weakness was caution: it accepted 82% of valid claims, and 60% on the beer game.1.000.950.900.850.750.90
llama3 8B (local, open weight)
With equations0.550.460.510.450.510.45
Reference
Generator (fidelity)Fidelity, not a competitor. The generator’s checks use the same rules as the labels, applied to recovered rather than exact equations, so this row shows how faithfully the recovered mechanism reproduces the true verdicts. It is not an advantage under equal information.1.000.991.000.991.001.00
Fable 5.1
Directional noteHow it was tested: the same 120-claim subset under the same restrictions as the other judges — every tool disabled, the same prompts, the same one-line verdict. On the simplest system it matched the other frontier models. On the beer game and Lorenz it often responded by writing a numerical-integration script to check the claim and, with no way to run it, ended without a verdict; nearly a quarter of its replies were unparsed, mostly for this reason. Directional only: the paper keeps it out of its tables because extra prompt engineering would have broken the head-to-head. It is not a measure of Fable’s capability, and with tool access it may perform strongly, which was not tested.Not scored. Run on the 120-claim subset: it matched the frontier judges on the tank, and on the hardest systems tried to write integration code it could not run.
* Separate run after the main benchmark, on a stratified 120-claim subset with the generator’s output only. Not directly comparable to the full-benchmark rows.
From the working paper. Tank to Lorenz runs from one linear state to chaotic dynamics. As judges, the frontier models hold up far better than they forecast, and best of all when handed the generator’s predicted trajectory. Tap the i beside a row for how it was tested.

Three things stood out.

Forecasting broke down; judging held up better. Sonnet and Haiku forecast the three simple systems well and fell apart on SEIR, the beer game and Lorenz. Sonnet produced a valid forecast for 2 of 20 SEIR scenarios and none for the beer game or Lorenz. Yet from a description alone, Sonnet judged claims about those same systems at 0.86 to 0.94 balanced accuracy (0.5 is chance). Verifying turned out to be easier than forecasting.

The generator's output was the most useful thing to hand a judge. With it, Sonnet held 0.94 to 1.00 at every level, and Haiku rose from 0.68 to 0.90 on the beer game and from 0.70 to 0.93 on Lorenz. Answer matching and hand-written rules stayed between 0.52 and 0.78. The small local model stayed near chance in every condition, because it accepted almost everything.

A local open-weight reasoning model came close. In a separate run on a 120-claim subset, gpt-oss 20B, given the generator's output, scored 0.89 against Haiku's 0.93 and Sonnet's 0.97, and caught 97% of the invalid claims. That run only tested the assisted condition, so it shows assisted performance, not how much the generator improved it. A separate run of Claude Fable 5.1 often responded to the hardest systems by writing integration code it couldn't run under our tool-free setup, and ended without a verdict; the paper reports it separately, not as a capability ranking.

These are constructed and model-generated claims, which may be easier to judge than mistakes in the wild. The paper covers the protocol, its two amendments, and every per-claim result.

Where generators break

A language model's characteristic failure is confident confabulation: a fluent answer that is wrong, and wrong differently from one run to the next. A deterministic generator repeats its verdict for the same inputs, but repeatability doesn't make it right. The paper highlights two failure modes: under-specification and overfitting.

Under-specification. A generator can test claims only about the mechanisms it models. Applied outside that scope, such as to a quality rating assembled from correlated signals, it may reject valid claims or give precise-looking verdicts without a sound basis. The fix is scope. Tie each check to a specific modeled mechanism and intervention, and report anything outside that scope as out of scope rather than scoring it.

Overfitting. A generator fitted to thin or narrow data can match its training runs and still miss real couplings, regimes or thresholds. It may give confident verdicts on training-like cases while failing on unfamiliar regimes. The fix is in how you build it. Use diverse interventional regimes, choose candidate terms from knowledge of the system rather than convenience, validate on held-out regimes including stressed ones, and re-validate when the data shifts.

Both come from the same root: treating recovered equations as the mechanism without checking that they are. A generator scoped to what it models, and validated on regimes it hasn't seen, is a strong evaluator. Trusted beyond that, it's a confident source of error.

Where they broke on us

Our benchmark showed why training-domain validation isn't enough. On SEIR and the beer game, recovery produced dense polynomials of 147 and 181 terms rather than the mechanism, yet both passed validation inside the training domain. On stressed regimes outside it, SEIR passed three of five tests and the beer game none, with errors up to seven times the tolerance. Validating only inside the training range would have accepted both.

Some misses were ours rather than the generators'. On Lorenz, a short periodic path passed every attractor check we'd written and was caught only by the flow rule. A rule check built from the exact equations scored 0.86 to 0.99 on the first five levels, then 0.47 on Lorenz, rejecting almost every valid claim. And in the fishery, one answer label was wrong because the simulated window ended before a slow collapse finished.

The guardrails belong around the evidence as well as the model. Validate on regimes the generator never saw, freeze scenarios and tolerances before scoring, disclose any changes, and don't treat two checks that share an input as independent confirmation.

Evaluating part of a flow

A generator doesn't have to evaluate a whole system at once. Interventional clamping fixes selected states to supplied or observed values and scores the rest over a chosen window. That opens up three practical uses.

The first is partial claims. A forecast that covers only inventory levels can be scored without modeling the whole supply chain: clamp the pipeline to its supplied or observed values and score the stock on its own. The second is localizing errors. When a full-system claim fails, clamping upstream states one at a time can help narrow down which component-level dynamics disagree with the claim.

The third is a possible extension: progressive evaluation during a run. At each observation, initialize the generator from the current state and compare proposed inputs over the next short window. Repeat as observations arrive, or target one stage or decision. Each update starts a new conditional question, and it must not overwrite an earlier counterfactual with later factual states affected by that intervention. This workflow was not tested in the benchmark.

I think of it as aim small, miss small: a short window anchored to real observations keeps each check focused and makes failures easier to locate. For a unit-specific counterfactual within a deterministic model, the starting state must determine future evolution under the proposed inputs, with no unresolved disturbances. Otherwise, hidden conditions require modeling and inference. Intervention-affected states must evolve freely, and it's worth reporting which states were clamped and over which window. The earlier benchmark scored 108 inventory claims over weeks 8 to 24 at 97% accuracy, comparing stock forecasts with a full generator rollout. That supports windowed scoring, not the full clamped counterfactual procedure.

Where to start

The economics are simple. A generator moves the cost of evaluation up front: interventional data, recovery and validation are the expensive part. The earlier three-system harness scored claims in milliseconds; Sonnet's median judgments in the ladder benchmark took 9 to 15 seconds. Passing generator trajectories to judges improved or preserved their point-estimate accuracy. Using the generator as a selective pre-filter, with its trajectory passed to a judge where you want a second opinion, is a deployment proposal, not a routing policy tested here. In AI Airlocks terms, it's semantic validation followed by optional model re-evaluation.

It isn't right for every problem. You need measurable state, rules that aren't only correlations, interventions you can observe, claims that map to state, and a mechanism stable enough to recover and re-validate. Where those hold, it's a strong complement to answer checks, LLM judges and human review, not a replacement for them.

Start with one system and one intervention. Validate it on regimes it hasn't seen, and extend its scope only as far as you understand the mechanism.


Where this sits in the Frontier series. The Foundation essays put governance into the architecture: replayable Steps and States, Causal Lineage, and Airlocks that validate before production. This essay looks at what that recorded context can support once it includes interventions and responses.

The paper. The full working paper, AI Eval against Causal Generators, has the definitions, protocol and both amendments, per-claim results and the benchmark harness, and is available on Zenodo with its Python release at doi:10.5281/zenodo.23060990.

Standing on. Pearl on causal models and interventions; Mooij and colleagues and Bongers and colleagues on causal models with feedback; Peters and colleagues on invariance and dynamical systems as causal models; Brunton, Proctor and Kutz for SINDy; Sterman's stock-and-flow tradition and his beer-game experiments; Schaefer and Clark for the harvested-fishery model behind the interactive example; and Zheng and colleagues on LLM-as-judge. The contribution here is putting these together and testing the result.