Methodology
Simulating a specific person is measurably closer to that person than a demographic average is.
Eight published results below, three objections that still stand, and one yardstick: a person re-asked two weeks later agrees with themselves 81.7% of the time.
00
The yardstick is not perfection. It is how consistent a person is with themselves.
Re-ask a person the same questions two weeks later and 81.7% of their answers match the ones they gave the first time.
Same instrument, same questions. The gap between 71.7 and 81.7 is what is left to close.
Simulated samples reproduced findings whose answers were already known
- Four classic studies, re-run with simulated participants instead of recruited ones.
- The known result came back in all four.
- A replication is the weakest test available, and the right place to start.
Four replications reported in the paper.
Fidelity comes from conditioning on a person, not on a demographic label
- Give the model one person's backstory and its answers follow that person.
- Give it a demographic cell and the answers converge on an average nobody holds.
Illustrative panel of five. No study data.
Simulated consumers price like consumers
- Asked what they would pay, simulated consumers slope downward.
- Willingness to pay falls as price rises, without being told that it should.
Illustrative demand curve. No study data.
An agent grounded in a person's own words predicts that person better than any profile of them
- An agent built from a person's own interview beats every profile written about that same person.
- Same person, four ways of describing them, four accuracies.
Share of the person's own two-week consistency, held-out GSS items.
One model predicts behaviour in experiments it was never tuned for
- One model, held out experiments, no per task fitting.
- Generalisation is the test a lookup table cannot pass.
Counts as reported in the paper.
Simulated respondents predict which message moves people, at the accuracy of pooled human forecasters
- Which of two messages moves people, predicted before the study runs.
- At the accuracy of a pool of human forecasters, not of one expert.
r = 0.85
Points illustrative. No study data. The correlation is the paper's.
Raw simulated effects overshoot, and calibration is what closes the gap
- Raw simulated effects come out larger than the real ones.
- Calibrated against known outcomes, the error falls by 77x.
Error relative to the calibrated fit.
08
Three published objections, and what we say back.
The three strongest results against the claim, including one from the group whose data we want to be judged on.
Persona prompting produces samples that are not representative of the group they name.
Bisbee, Clinton, Dorff, Kenkel and Larson, 2024 · Political AnalysisWe agree, and it is the same finding as entry 02. A demographic persona is the condition Park et al. measured at 74 percent, the weakest arm they tested. Audiences here are built from records of individuals, and a distribution is not reported until it has been scored against a real one.
Underspecified prompts collapse onto the dominant view of whichever group is named.
Santurkar et al., 2023 · ICML, OpinionQAUnderspecification is the failure mode, so specification is the method. A panel built from what people actually said carries their disagreement with it, and an audience definition that cannot be resolved to individuals is refused rather than averaged.
On genuinely novel outcomes, correlation with real effects drops to r = 0.20.
Peng et al., 2025 · arXiv:2509.19088, from the Twin-2K-500 groupThis is the binding constraint, and it comes from the group whose data we want to be judged on. Recalling a held-out answer and predicting a novel outcome are different problems with different accuracy; it is why entry 07 matters, why calibration against known outcomes precedes any forecast, and why our own sets stay sealed until they are run.







