← evacueedoes it port? → wildfire

What the model learns — and where it diverges from real people

We run a population of LLM residents through the real 2024 Noto tsunami, read each one's path-dependent decisions, and invert that behaviour into a set-identified trait → factor → decision structure of the LLM-agent model. Every figure below is a real run.

Honesty spine. We recover the factor structure of an LLM-agent model, not a validated causal mechanism in humans. Stability and replication are evidence about the model. Our only contact with human ground truth is the held-out odds-ratio test — out-of-sample corroboration, not validation. So the deliverable is a set of admissible structures Θ, and we report its size.
the deliverable

A set-identified trait → factor → decision structure

The backbone the data forces (stable edges shared by every admissible graph): which persona traits, through which latent behavioural factors, drive which decisions. We do not claim one true graph — we report the admissible set Θ and its size.

trait factor decision causal structure
the honest result

Held-out reality test: 0 / 0 corroborated

The real survey's odds-ratios were held out — never used to fit. The model reproduces the common-sense effect (anticipating a tsunami → evacuate) but FLIPS the sign on the counterintuitive human paradoxes: in reality prepared people (emergency bag) evacuate LESS and a family plan reduces car use — the LLM does the opposite. The held-out test is what surfaces this.

This is the payoff of the honesty spine: a do(world)-robust model structure is not human truth, and the one ground-truth contact tells us exactly where the LLM's prior misses the human paradox.

causal orientation

do(trait): undefined / undefined edges oriented within the model

To turn associations into oriented edges we intervene: flip one trait, hold everything else fixed, re-ask the LLM, and measure how the factor's activation moves. An edge that the flip moves is causal within the model — and removes its orientation ambiguity from Θ. (An intervention on the model, not on people.)

do(trait) interventions
the closed loop

Refiner feedback sharpened undefined / 3 load-bearing edges

The loop the design is built around: the structure ranks which traits are load-bearing; the refiner gives exactly those trait dimensions a richer first-person narrative; we re-run; the edges become more identifiable (a larger carrier-vs-noncarrier activation gap).

the automatic finder

Diversity-driven saturation + cross-world portability

Not answer-directed: we just draw large diverse samples until the discovery curve flattens, then re-run in a structurally different world (a faster tsunami). Factors that recur in both worlds are portable. The car-share gap does NOT close with more sampling — a stable LLM bias, calibration's job, not coverage's.

finder discovery, convergence, portability
the simulator side

A small basis of populations; the rest are mixtures

Each admitted cohort is a point in indicator space. Most are convex mixtures of a few canonical ones — the basis. The real survey target sits just outside the hull on the car axis (the same LLM bias).

outer loop basis + factor replication + discovery
the engine must grow

A rule-based ABM holds only —% of what residents say

Residents say more than 'go to the nearest shelter by foot or car' — they pick up a child, wait-and-see, follow neighbours. As we grow the action grammar by the most-demanded construct, coverage climbs to ~99% and flattens. The uncovered residual is exactly what feeds factor discovery.

growing grammar coverage
how much is enough

Scale & fidelity ablation

Below ~200 samples the surrogate overfits and zero factors are discoverable; sample size dominates, physics-tick resolution barely matters once you have enough samples.

N x fidelity ablation grid
N x physics resolution grid
one resident, two paths

The cognitive × physical hybrid trace

One resident's physical path on the real road graph, time-aligned to their cognitive path (felt-risk and the factor cited at each step), aggregating into the population causal graph. Factor labels via a frozen extractor, in the lineage of think-aloud coding and Bayesian inverse planning.

hybrid cognitive physical map