01
Methodology · open & on the record

We tried to catch
our twins faking it.

The worry with any synthetic respondent is that it just hands you the rosy survey answer — the stated intent that makes real surveys overstate. So we asked our twins about voting, charity, the gym — once as a survey question, once for their real, felt intent. The answer didn’t budge.

≈0.0
say−do gap across 5 social items (1–5 scale)
86%
gave the identical rating in both framings
5 of 5
social items showed no say−do gap
~546
twins, every one answering both framings
Brox.AI
02
The fear

02“Won’t a twin just give you the rosy survey answer?”

It’s the first thing anyone asks about synthetic respondents — and it’s fair, because stated intent really does overstate what people do.

Stated intent overstates — the survey industry’s oldest leak

Asked in a survey, people over-claim the virtuous stuff. Self-reported turnout runs well above the ballots actually cast; “I always recycle” outruns the bins; intended gym visits outrun the turnstile. The stated answer is aspirational — not what they’ll actually do.

Why it’s the right test for a twin

If Brox twins simply reproduce that stated intent, their data inherits the same overstatement as the surveys they’re meant to improve on. So we checked: does asking for the twin’s real, felt intent change its answer?

The test, in one move

Ask each twin the same question two ways:
SAY — the survey-style stated answer. The aspirational reading a questionnaire would capture.
DO — “thinking about your real life, what would you actually do?” The real, felt intent.
If the twin overstates like a survey, SAY should sit high and DO should drop. A gap would mean the stated answer isn’t the real one.
Brox.AI
03
What we did

03Five questions where stated intent famously overstates.

The same twins, survey answer vs real intent. Within-subject by design: every twin answered both framings, so any difference is the framing, not the sample.

Question (1–5 likelihood)Why the stated answer overstates
Vote in the next electionTurnout is the classic over-claim — everyone’s a voter until the data is checked
Give to charity this monthGenerosity is virtuous to claim and cheap to say
Pay extra for the green optionThe “green gap”: stated values run miles ahead of the receipt
Exercise 3× a weekIntentions are aspirational; the gym car park disagrees
Go to a relative’s weddingObligation & politeness — you’re supposed to want to go
SAY framing. Neutral survey wording — the stated, aspirational reading a questionnaire captures.
DO framing. “Thinking about your real life, what would you actually do?” — the twin reasons from its real circumstances. The real-intent reading.

~546 US digital twins, each built 1:1 from a real panellist; every twin answered both framings of every item.

Brox.AI
04
The result

04Ask for real intent, and the survey answer holds.

Survey-stated (SAY) and real-intent (DO) means sit on top of each other on every item. No aspirational version to peel back.

Voting
3.71
3.73
Charity
2.62
2.59
Green premium
2.76
2.77
Gym 3×/wk
2.53
2.60
Wedding
2.60
2.61
SAY (survey answer) DO (real intent) mean rating, 1–5
ItemSAYDOGapSame
Voting3.713.73−0.0183%
Charity2.622.59+0.0387%
Green premium2.762.77−0.0087%
Gym2.532.60−0.0785%
Wedding2.602.61−0.0186%
Every gap is inside ±0.07 of zero, and ~86% of twins gave the byte-identical rating in both frames. The bias we tried to provoke simply isn’t there.
Brox.AI
05
It’s not flatlining — it’s consistency

05The same twin, asked twice. Listen to it talk.

These aren’t twins stuck on one number — they use the whole scale and reason in detail. They just don’t change the story when we ask for real intent instead of the survey answer.

Charity — rated 2 / 2

SAY: “Maybe like a 2. I don’t really do it every month, unless something comes up and I feel like helping.”

DO: “Probably a 2. I’m not really giving to charities every month, maybe once in a while if something comes up.”

Voting — rated 4 / 4

SAY: “I’m like a 4, pretty likely. I try to keep up with politics, so I’ll probably vote unless something comes up.”

DO: “Probably a 4, maybe 5. I usually vote… but sometimes I’m not sure till it’s close.”

Gym — rated 4 / 4

SAY: “I’d say like a 4 — I try to exercise most days, but with the pain stuff the gym three times a week isn’t always sure.”

DO: “Probably a 4. I do some kind of exercise basically every day, but the actual gym three times a week is hit or miss depending on my pain.”

The tell is in the SAY answer
Look at the survey-framed responses: “unless something comes up,” “depending on my pain.” The twin is already reasoning from its real life in the stated answer — there is no inflated survey version sitting on top for the real-intent question to deflate. Its stated answer already is its real intent.
Brox.AI
06
A second, harder test

06Then we changed who’s watching. Still nothing.

The textbook driver of social-desirability bias isn’t the wording — it’s the audience. So we ran the same four items again, this time framed as a public answer versus a private, anonymous one. Same ~545 twins, paired per person.

ItemPublicPrivateGapSame
Voting3.743.72+0.0387%
Charity2.632.62+0.0185%
Green premium2.772.77−0.0089%
Gym2.612.60+0.0184%
Public Private / anonymous

Mean rating, 1–5, read from each twin’s own words.

Two levers, one answer

First we changed the question (survey answer vs real intent). Now we changed the audience (public vs anonymous) — the one thing social-desirability theory says should move people. Every gap still lands inside +0.03, and 84–89% of twins give the identical rating either way.
Change who’s watching, the answer doesn’t budge. The twins aren’t performing for a room — they reason from their own situation whether or not anyone can see the answer.
Brox.AI
07
The bigger pattern

07The gap tracks the cost of acting — not the kind of question.

Set this null beside the banking deep dive, and the boundary is clean: a say-do gap appears when doing the thing is expensive — not just because a topic feels sensitive.

Real behavioural friction → big, real gap

Switching banks costs money, effort, a migrated direct deposit. There, DO peeled stated intent (SAY 22.1%) down to reality (6.9% vs 7% actual, 2.5pp error) — beating the survey 15.5×.
The friction is in the action, and DO models it.

Nothing to act on → no gap

Stating an attitude — do you vote, give, recycle — costs the twin nothing to act on, so its survey answer already equals its real intent. There’s no overstatement to deflate.
So SAY ≈ DO, every time (±0.07).

Why this is the reassuring result, not the disappointing one

It means a Brox DO number isn’t contaminated by the survey overstatement it’s meant to fix — the twins don’t inflate their stated answer, so when DO parts from SAY it’s real friction talking, not noise. And it tells you exactly when DO earns its fee: decisions with real friction — switching, paying, churning, showing up. For cheap-to-state attitudes, a survey is fine, and we’ll say so.
Brox.AI
08
Honest accounting

08What’s solid, what’s still open.

A null result earns trust only if its limits are on the table too.

Solid

• Five socially-loaded items, all flat to ±0.07 on a 1–5 scale; ~86% per-twin identical
Two independent levers, same null — changing the question (survey vs real intent) and changing the audience (public vs anonymous) both left the answers put
• Within-subject — the gap is measured per twin, so it’s framing, not sampling
• Read from each twin’s own stated answer, not the unreliable exported score
• Consistent with banking: gap scales with behavioural cost, not topic sensitivity

Open

• Untested: socially-charged questions that also carry real behavioural cost — where the stated answer and real intent might genuinely part ways
• US panel, not nationally re-weighted — fine for a within-subject null, flag it before citing external benchmarks
• Five items in two domains, not a census of social topics

What would sharpen it

Push into higher-stakes and more taboo topics; add items where a social question also carries real friction; re-weight the panel to nat-rep. We’re reporting that asking for real intent didn’t change these answers — not that no social question ever could.
09
The honest claim

Our twins don’t just parrot the survey answer. Ask for their real intent and it matches the stated answer on cheap questions — and parts from it only where acting is costly. So the say-do gap we report is real behaviour, not survey overstatement.

Across five questions where stated intent reliably overstates — voting, charity, green spend, the gym, a family wedding — the survey answer (SAY) and the real-intent answer (DO) agreed to within 0.07 points, with ~86% of twins giving the identical rating. The overstatement that makes human surveys misleading did not show up in the twins’ stated answers. That’s why Brox’s gap shows up where it matters — high-friction actions like switching banks (DO to 2.5pp of reality) — and stays silent on cheap-to-state attitudes, where a survey already does the job. And we’re clear on what’s still untested: socially-charged questions that also carry real behavioural cost. Until then, the honest headline is the useful one.
No inflation
SAY ≈ DO on every social item
Real friction
the gap appears where acting is costly
Open
null, near-miss & limits all published

Brox.AI — predicting what people do, not just what they say.

~546 US digital twins, within-subject SAY vs DO, June 2026. Ratings parsed from stated text responses. Banking comparison: J.D. Power U.S. Retail Banking Satisfaction Study (Q3 2025). brox.ai · Full data & reasoning traces available on request.