← the research

Does it really work? Porting the whole method to a wildfire

The pipeline was built on the 2024 Noto tsunami. The real test of a method is whether it transfers. So we pointed the same machinery — persona sampling → path-dependent LLM decision traces → factor discovery → a set-identified trait → factor → decision structure → a held-out reality test — at a structurally different disaster: the 2018 Camp Fire (Paradise, CA). Nothing was reused but the method; the scenario was built from the wildfire-evacuation literature.

what we set out to explore

A hazard that breaks the tsunami's every assumption

A wildfire is a moving, wind-driven front, not a single impulse. Escape is horizontal along a few congestible roads, not vertical to high ground. The dominant mode is the private car (>94%), not foot. The official alert failed — only ~19% ever got an evacuation order; ~45% first learned by seeing the fire themselves. And there is a stay-and-defend option with no tsunami analogue. The question: does the method still run, and does it discover the right, different factors?

what we got — 1 · it ran

The method ported end-to-end and discovered different, hazard-appropriate factors

No tsunami factors carried over. From the wildfire traces the finder surfaced mechanisms specific to fire: a stay-and-defend self-efficacy, sensory confirmation (act on seeing flames, not on a warning), dependent/pet attachment, authority dependence, and a fire past-experience bias. The method adapts; it does not regurgitate Noto.

wildfire trait-factor-decision structure
what we got — 2 · the honest reality test (and a rigor fix)

Held-out reality test

Each held-out effect is tested against the variable its source actually measured: rate effects (evacuate vs stay) as an odds ratio on evacuating, timing effects as the shift in simulated departure minute among those who left. This matters. An earlier draft wrongly scored "leaves later" as "evacuates less" and produced false sign-flips for longtime-residents and the elderly. Corrected, all five timing directions reproduce — longtime-residents (Δ+11 min) and the elderly leave LATER; smartphone/plan/seeing-fire leave EARLIER. The two genuine flips are rate paradoxes the LLM misses: prior fire awareness should LOWER evacuation (habituation), but the model raises it — the wildfire echo of Noto's emergency-bag paradox. The strongest result: believes-can-defend → does NOT evacuate (OR 0.07), the stay-and-defend mechanism, corroborated.

Outcome-matched: each effect is tested against the variable its source actually measured — rate effects (evacuate vs stay) as an odds ratio on evacuating, timing effects as the shift in simulated departure minute among those who left. Targets are literature-composite (Grajdura & Niemeier 2022; McCaffrey 2018; Marshall; Kincade) — a directional-consistency check, NOT a single-survey mechanism test like Noto's Hasegawa table.

the cross-hazard finding

The same method, the same honest limit — on two hazards

Noto tsunami
held-out OR corroborated: 1/5
flips the bag paradox (prepared → evacuates less)
Camp Fire wildfire
right sign: 8/10 · strict: 2/10
flips on the rate paradox (prior fire awareness → should lower evacuation)

The method is portable — the same machinery ran on a hazard that breaks every tsunami assumption and surfaced the right, different factors. And the honesty spine holds across hazards: a do(world)-robust model structure is not human truth, and the one ground-truth contact (held-out indicators) shows the LLM captures the intuitive levers but misses the counterintuitive human structure — on both disasters.