Methodology

Simulating a specific person is measurably closer to that person than a demographic average is.

Eight published results below, three objections that still stand, and one yardstick: a person re-asked two weeks later agrees with themselves 81.7% of the time.

08published results
03objections, answered
81.7the human ceiling
77xcalibration error reduction

Eight papers

select a cover to open its first two pages in its row

00

The yardstick is not perfection. It is how consistent a person is with themselves.

Re-ask a person the same questions two weeks later and 81.7% of their answers match the ones they gave the first time.

59.2chance
71.7their digital twin · 87.7% of the ceiling
81.7the same person, later
81.7the same person, two weeks later
71.7their digital twin · 87.7% of the ceiling
59.2chance

Same instrument, same questions. The gap between 71.7 and 81.7 is what is left to close.

Toubia et al., 2025 · arXiv:2505.17479Read the paper
012022

Simulated samples reproduced findings whose answers were already known

  • Four classic studies, re-run with simulated participants instead of recruited ones.
  • The known result came back in all four.
  • A replication is the weakest test available, and the right place to start.
Aher, Arriaga and Kalai, 2022 · arXiv:2208.10264Read the paper
Replications
Ultimatum gamereplicated
Garden path sentencesreplicated
Milgram shockreplicated
Wisdom of crowdsreplicated

Four replications reported in the paper.

022023

Fidelity comes from conditioning on a person, not on a demographic label

  • Give the model one person's backstory and its answers follow that person.
  • Give it a demographic cell and the answers converge on an average nobody holds.
Argyle et al., 2023 · arXiv:2209.06899Read the paper
Five people, five spreads
GROUP AVERAGEP1P2P3P4P5

Illustrative panel of five. No study data.

032023

Simulated consumers price like consumers

  • Asked what they would pay, simulated consumers slope downward.
  • Willingness to pay falls as price rises, without being told that it should.
Brand, Israeli and Ngwe, 2023 · HBS Working Paper 23-062Read the paper
Demand curve
QUANTITYPRICE

Illustrative demand curve. No study data.

042024

An agent grounded in a person's own words predicts that person better than any profile of them

  • An agent built from a person's own interview beats every profile written about that same person.
  • Same person, four ways of describing them, four accuracies.
Park et al., 2024 · arXiv:2411.10109Read the paper
Predictive accuracy
interview + survey86
interview only83
survey only82
demographics74

Share of the person's own two-week consistency, held-out GSS items.

052025

One model predicts behaviour in experiments it was never tuned for

  • One model, held out experiments, no per task fitting.
  • Generalisation is the test a lookup table cannot pass.
Binz et al., 2025 · arXiv:2410.20268Read the paper
Scale of the test
160experiments
60k+people
10M+choices

Counts as reported in the paper.

062026

Simulated respondents predict which message moves people, at the accuracy of pooled human forecasters

  • Which of two messages moves people, predicted before the study runs.
  • At the accuracy of a pool of human forecasters, not of one expert.
Ashokkumar, Hewitt, Ghezae and Willer, 2026 · Nature 656, 115-122Read the paper
Predicted vs observed
OBSERVED EFFECTPREDICTED EFFECT

r = 0.85

Points illustrative. No study data. The correlation is the paper's.

072026

Raw simulated effects overshoot, and calibration is what closes the gap

  • Raw simulated effects come out larger than the real ones.
  • Calibrated against known outcomes, the error falls by 77x.
Hut and Masoero, 2026 · arXiv:2608.02345Read the paper
Error, raw vs calibrated
raw77x
calibrated1x

Error relative to the calibrated fit.

08

Three published objections, and what we say back.

The three strongest results against the claim, including one from the group whose data we want to be judged on.

Objection 01

Persona prompting produces samples that are not representative of the group they name.

Bisbee, Clinton, Dorff, Kenkel and Larson, 2024 · Political Analysis
Our response

We agree, and it is the same finding as entry 02. A demographic persona is the condition Park et al. measured at 74 percent, the weakest arm they tested. Audiences here are built from records of individuals, and a distribution is not reported until it has been scored against a real one.

Objection 02

Underspecified prompts collapse onto the dominant view of whichever group is named.

Santurkar et al., 2023 · ICML, OpinionQA
Our response

Underspecification is the failure mode, so specification is the method. A panel built from what people actually said carries their disagreement with it, and an audience definition that cannot be resolved to individuals is refused rather than averaged.

Objection 03

On genuinely novel outcomes, correlation with real effects drops to r = 0.20.

Peng et al., 2025 · arXiv:2509.19088, from the Twin-2K-500 group
Our response

This is the binding constraint, and it comes from the group whose data we want to be judged on. Recalling a held-out answer and predicting a novel outcome are different problems with different accuracy; it is why entry 07 matters, why calibration against known outcomes precedes any forecast, and why our own sets stay sealed until they are run.