What the model learns — and where it diverges from real people
We run a population of LLM residents through the real 2024 Noto tsunami, read each one's path-dependent decisions, and invert that behaviour into a set-identified trait → factor → decision structure of the LLM-agent model. Every figure below is a real run.
A set-identified trait → factor → decision structure
The backbone the data forces (stable edges shared by every admissible graph): which persona traits, through which latent behavioural factors, drive which decisions. We do not claim one true graph — we report the admissible set Θ and its size.
Held-out reality test: 0 / 0 corroborated
The real survey's odds-ratios were held out — never used to fit. The model reproduces the common-sense effect (anticipating a tsunami → evacuate) but FLIPS the sign on the counterintuitive human paradoxes: in reality prepared people (emergency bag) evacuate LESS and a family plan reduces car use — the LLM does the opposite. The held-out test is what surfaces this.
This is the payoff of the honesty spine: a do(world)-robust model structure is not human truth, and the one ground-truth contact tells us exactly where the LLM's prior misses the human paradox.
do(trait): undefined / undefined edges oriented within the model
To turn associations into oriented edges we intervene: flip one trait, hold everything else fixed, re-ask the LLM, and measure how the factor's activation moves. An edge that the flip moves is causal within the model — and removes its orientation ambiguity from Θ. (An intervention on the model, not on people.)
Refiner feedback sharpened undefined / 3 load-bearing edges
The loop the design is built around: the structure ranks which traits are load-bearing; the refiner gives exactly those trait dimensions a richer first-person narrative; we re-run; the edges become more identifiable (a larger carrier-vs-noncarrier activation gap).
Diversity-driven saturation + cross-world portability
Not answer-directed: we just draw large diverse samples until the discovery curve flattens, then re-run in a structurally different world (a faster tsunami). Factors that recur in both worlds are portable. The car-share gap does NOT close with more sampling — a stable LLM bias, calibration's job, not coverage's.
A small basis of populations; the rest are mixtures
Each admitted cohort is a point in indicator space. Most are convex mixtures of a few canonical ones — the basis. The real survey target sits just outside the hull on the car axis (the same LLM bias).
A rule-based ABM holds only —% of what residents say
Residents say more than 'go to the nearest shelter by foot or car' — they pick up a child, wait-and-see, follow neighbours. As we grow the action grammar by the most-demanded construct, coverage climbs to ~99% and flattens. The uncovered residual is exactly what feeds factor discovery.
Scale & fidelity ablation
Below ~200 samples the surrogate overfits and zero factors are discoverable; sample size dominates, physics-tick resolution barely matters once you have enough samples.
The cognitive × physical hybrid trace
One resident's physical path on the real road graph, time-aligned to their cognitive path (felt-risk and the factor cited at each step), aggregating into the population causal graph. Factor labels via a frozen extractor, in the lineage of think-aloud coding and Bayesian inverse planning.







