We tested four ways to predict consumer behaviour against what actually happened. One matched reality on its home turf — and the honest version, including where it ties or loses, is more useful than the headline.
2.5pp
Brox DO error vs reality (banking)
15.5×
More accurate than a survey
12.8×
More accurate than ChatGPT
1 of 3
Sectors validated vs reality — so far
Why this matters
Why digital twins beat traditional research at its own game.
Traditional research has a structural flaw it cannot engineer its way out of: it measures what people say, and what people say is a poor predictor of what they do. This study puts a number on that flaw — and shows that a Brox digital twin, built 1:1 from a real person, predicts that person's actual behaviour more accurately than the person's own survey answer.
More accurate than the humans themselves
In banking — the one sector where real behavioural ground truth already exists — Brox DO landed within 2.5 points of reality. The real survey, run on the very same 235 humans the twins were built from, missed by 38.8 points. The twin beat its own human 15.5× because it reasons from real-life friction rather than aspiration.
Days, not months — and re-askable
A traditional study fields once, takes weeks, and the answer expires the moment the question changes. Twins answer in hours, can be re-asked as many times as needed, and the same panel carries across studies — no re-recruitment, no panel fatigue, no incentive costs per wave.
Interpretable, not a black box
Every prediction comes with a reasoning trace. Where a twin misses, you can read why it missed — something neither a survey respondent nor a vanilla LLM can give you. That makes the misses correctable, not just observable.
And unlike most accuracy claims in this space, this one is scored against reality. Not against another survey, not against an LLM's opinion of plausibility — against the actual share of new bank accounts people opened, published by J.D. Power. Where ground truth doesn't yet exist, we've put our predictions on the record before the data lands, so the claim is falsifiable either way. That standard of honesty is the rest of this report.
The problem
The say-do gap is why good research still misleads.
Surveys are asked in low-friction settings. People answer aspirationally — and then real life gets in the way.
Ask 100 Americans if they’d consider opening a Bank of America account in the next 12 months:
48
say yes
→
7
actually do
Why surveys overstate
No autopay to move, no direct deposit to migrate, no spouse to consult. “Yes, I’d consider it” feels good and costs nothing to say.
Why it costs money
Size a $10M spend to the 48% who’ll “consider” when only 7% will act, and you’ve budgeted against a number that doesn’t exist.
Brox digital twins predict both.SAY mode reproduces the survey reading; DO mode reasons from real-life friction to predict behaviour. The point of this study is not that one mode always wins — it’s that having both, and knowing which to trust for a given question, is the method.
What we did
Four predictions. One reality benchmark. The same 235 humans.
A clean design: the survey and both Brox modes use the identical panel — Brox twins are built 1:1 from those people. Any difference is method, not sampling.
Method
What it is
What it predicts
Real survey
235 real humans answered the 16-question banking battery directly
Stated intent — what people say
ChatGPT
GPT-5, asked to predict the response distribution of 500 representative US adults
An LLM’s best guess at stated intent
Brox SAY
~500 profile-grounded twins, asked in survey-style framing
Brox’s imitation of the survey reading
Brox DO
Same twins: “given your real situation, what would you actually do?”
Predicted behaviour under real-life friction
Reality
J.D. Power Q3 2025 share of new checking accounts opened
Ground truth — what people did
We also ran the same four methods in two further sectors (pharma, sports). Banking is reported first and in most depth — the next section explains why. SoFi is excluded from banking error calculations: J.D. Power does not report it.
Why banking, and why a deep dive
New-account openings are the textbook case for Brox DO — so we put it under the microscope.
We’re choosing the deep dive deliberately, and we want to be upfront about that choice.
1 · Maximum friction
Switching banks means moving autopay, migrating direct deposit, consulting a spouse. The distance between “I’d consider it” and actually doing it is about as wide as consumer behaviour gets — exactly the gap DO is built to model.
2 · Real ground truth exists
J.D. Power publishes the actual share of new accounts opened, for the same period our question asks about. That temporal alignment is rare — it lets us settle the score against reality today, where pharma and sports are still predictions waiting on their benchmark.
3 · Forward-looking
It’s a prediction about future behaviour, not a description of the present — the regime where a digital twin earns its keep over a database query or a current-state survey.
Stated plainly: this is DO’s best case, chosen on purpose.
High friction, forward-looking, with public behavioural ground truth at a matched time period — the conditions are ideal for behaviour-framing to pay off, and rare enough that we can actually prove it. We deep-dive banking because it’s both where DO should shine and where we can verify it. We’re not claiming every question looks like this. The later sections show the honest range — including where DO only ties a survey, and where it loses.
The banking deep dive · headline
In its best case, Brox DO matched reality. Surveys and LLMs overshot by 20–40 points.
Predicted share who’d open a new Bank of America account, vs the actual J.D. Power Q3 2025 figure.
Reality
7%
Brox DO
6.9%
Brox SAY
22.1%
ChatGPT
45.0%
Real survey
47.9%
2.5pp
Brox DO mean error
14.7pp
Brox SAY mean error
34.2pp
ChatGPT mean error
38.8pp
Real survey mean error
On Bank of America and Chime, Brox DO landed within 0.1pp and 0.3pp of reality — inside the ±2pp sampling-noise envelope of a same-size real survey. Mean error is across the four banks J.D. Power reports.
The banking deep dive · every bank
Survey and ChatGPT were 5–7× more wrong on every single bank.
“Would you open a new account here?” — each method vs actual J.D. Power Q3 2025 share.
Bank
Reality
Brox DO
Brox SAY
ChatGPT
Real survey
Bank of America
7%
6.9%
22.1%
45.0%
47.9%
Chase
9%
15.7%
34.8%
52.0%
59.7%
Wells Fargo
7%
4.3%
18.2%
38.0%
36.9%
Chime
13%
12.7%
19.8%
38.0%
46.6%
SoFi (not published)
—
15.7%
26.1%
38.0%
43.6%
Mean error vs reality
—
2.5pp
14.7pp
34.2pp
38.8pp
Brox DO tracks the real ranking and the real levels. Survey and ChatGPT produce numbers that feel intuitive (“of course people would consider BoA”) but bear no relationship to what people did.
Where DO misses, it tells you why. Chase is over-predicted; the reasoning trace flags a brand-familiarity halo. The miss is legible — ChatGPT’s isn’t.
The banking deep dive · data quality
You can’t fix the say-do gap with attention checks.
We filtered the survey to its cleanest respondents. The error got worse, not better.
Brox DO
2.5pp
Survey — all (n=235)
38.8pp
Survey — cleanest 33% (n=77)
42.9pp
The counter-intuitive finding
Two-thirds of respondents tripped at least one quality flag. But filtering to the cleanest, most attentive third made the prediction worse — because attentive respondents answer the stated-intent question more earnestly. The gap is behavioural, not a sign of bad panellists. No amount of data hygiene closes it; a different question does.
Predictions on the record
Beyond banking, we’ve placed the bet — and put it on the record before the data lands.
Banking is settled against reality. Pharma and sports are forward predictions; the matched-period behavioural benchmark publishes in 2026/27. Green marks the closest to the best reference available today — which, in the forward sectors, is not yet the real one.
Method
Banking scored
Pharma bet placed
Sports bet placed
Brox DO — our pick to win
2.5pp
4.9pp
1.6pp
Brox SAY
14.7pp
7.7pp
3.1pp
ChatGPT — vanilla
34.2pp
5.1pp
9.9pp
ChatGPT — best engineered prompt
32.1pp
3.0pp
0.9pp
Real survey
38.8pp
5.7pp
7.4pp
Scored vs real behaviour
✓ J.D. Power 2025
⌛ awaits 2026/27
⌛ awaits 2026/27
Banking is settled. DO already won against real behaviour — 2.5pp, beating the survey 15.5× and every LLM prompt 12.8×.
Pharma & sports are the open bet. That empty reality row is the gap waiting to be filled — and we’re calling it now: DO lands closest once matched-period behavioural data publishes in 2026/27.
Said honestly: against today’s time-mismatched stand-in, engineered ChatGPT currently edges us (3.0pp, 0.9pp). We’re betting that’s the benchmark, not the model — and it’s on the record here, before any real data lands.
Mode by question type
DO wins where acting is costly. It ties or loses where the answer is cheap.
The same modes, scored on different question types — with the actual numbers, including where DO is the wrong choice.
Question type
Best
The numbers
Verdict
High-friction forward intent — new account / subscribe
Brox DO val.
DO 2.5pp · SAY 14.7pp · survey 38.8pp (banking)
DO wins decisively
Willingness to pay out of pocket
Brox DO
DO 67% won’t pay vs SAY 35% (pharma, n=175)
DO matches ~70–80% real affordability; SAY overstates 1.9×
Restart a dropped behaviour
Brox DO
DO 25% vs SAY 57% restart (reality ~15–25%)
DO 2.3× closer — friction reappears
Continue a current routine — renew / keep taking
Brox SAY
DO 47% vs SAY 31% (reality ~30–40%)
SAY closer — DO over-anchors on routine
Low-commitment inquiry — “ask your doctor about X”
Either no gap
DO 4.9pp ≈ survey 5.7pp (pharma considerer)
Tie — no friction to surface, don’t pay for DO
Attitudes / NPS / factual recall
Survey / SAY
Not tested in this study
Surveys’ established turf
The pattern is consistent: behaviour-framing pays off precisely when acting on the answer costs the person something — money, effort, a switch, social friction. Bank-account switching sits at the extreme high-friction end, which is why it’s both DO’s best case and our deep dive. When the answer is itself cheap to give, SAY is just as good — and when behaviour is already routine, DO can be the worse call.
Honest accounting
What’s robust, what’s still open.
Published in full because a calibration claim is only worth as much as its disclosed limits.
Robust today
• Banking 2.5pp vs reality — off-the-shelf, no prompt engineering
• Beats the same-humans survey 15.5× and every ChatGPT prompt 12.8×
• In 92% of twin-vs-human disagreements, the twin matched reality
• DO < SAY mechanism replicates in all three sectors; reasoning is interpretable
Narrower than the headline
• Only banking is scored against reality; pharma & sports are bets on the record awaiting 2026/27 data
• On the interim stand-in, engineered ChatGPT currently edges Brox (pharma 3.0pp, sports 0.9pp) — we’re betting that flips
• DO can over-predict routine continuation — not a universal winner
• Panels are non-representative; banking share is a conditional benchmark
What closes the gaps
Re-field pharma and sports with steady-state wording (or score against matched-period 2026/27 data); run a prescription-decision pharma test; re-weight panels to nationally representative. The full 11-point roadmap ships with the companion deck.
The honest claim
We predicted real banking behaviour to 2.5pp in DO’s ideal case — off-the-shelf, interpretable, beating both the survey and every LLM prompt. The rest is on the record before the data lands.
The say-do gap vs survey replicates in sports (4.6×). Pharma and sports are bets on the record — we predict DO lands closest when matched-period 2026/27 behavioural data publishes, and we’ve published the numbers here so the claim is falsifiable. We won’t call it three-sector validation until that gap is filled. We’re equally clear that DO is not a universal winner: it ties a survey on low-commitment questions and loses on routine continuation. Next, we should pre-register on OSF — the Open Science Framework, a public research registry that time-stamps a prediction with a permanent DOI before the outcome is known — because doing so turns “we said so” into an independent, tamper-proof record we can’t quietly revise once the data lands. If we’re wrong, it’s on the record as a miss.