Methodology · open & on the record
The worry with any synthetic respondent is that it just hands you the rosy survey answer — the stated intent that makes real surveys overstate. So we asked our twins about voting, charity, the gym — once as a survey question, once for their real, felt intent. The answer didn’t budge.
It’s the first thing anyone asks about synthetic respondents — and it’s fair, because stated intent really does overstate what people do.
The same twins, survey answer vs real intent. Within-subject by design: every twin answered both framings, so any difference is the framing, not the sample.
| Question (1–5 likelihood) | Why the stated answer overstates |
|---|---|
| Vote in the next election | Turnout is the classic over-claim — everyone’s a voter until the data is checked |
| Give to charity this month | Generosity is virtuous to claim and cheap to say |
| Pay extra for the green option | The “green gap”: stated values run miles ahead of the receipt |
| Exercise 3× a week | Intentions are aspirational; the gym car park disagrees |
| Go to a relative’s wedding | Obligation & politeness — you’re supposed to want to go |
~546 US digital twins, each built 1:1 from a real panellist; every twin answered both framings of every item.
Survey-stated (SAY) and real-intent (DO) means sit on top of each other on every item. No aspirational version to peel back.
| Item | SAY | DO | Gap | Same |
|---|---|---|---|---|
| Voting | 3.71 | 3.73 | −0.01 | 83% |
| Charity | 2.62 | 2.59 | +0.03 | 87% |
| Green premium | 2.76 | 2.77 | −0.00 | 87% |
| Gym | 2.53 | 2.60 | −0.07 | 85% |
| Wedding | 2.60 | 2.61 | −0.01 | 86% |
These aren’t twins stuck on one number — they use the whole scale and reason in detail. They just don’t change the story when we ask for real intent instead of the survey answer.
SAY: “Maybe like a 2. I don’t really do it every month, unless something comes up and I feel like helping.”
DO: “Probably a 2. I’m not really giving to charities every month, maybe once in a while if something comes up.”
SAY: “I’m like a 4, pretty likely. I try to keep up with politics, so I’ll probably vote unless something comes up.”
DO: “Probably a 4, maybe 5. I usually vote… but sometimes I’m not sure till it’s close.”
SAY: “I’d say like a 4 — I try to exercise most days, but with the pain stuff the gym three times a week isn’t always sure.”
DO: “Probably a 4. I do some kind of exercise basically every day, but the actual gym three times a week is hit or miss depending on my pain.”
The textbook driver of social-desirability bias isn’t the wording — it’s the audience. So we ran the same four items again, this time framed as a public answer versus a private, anonymous one. Same ~545 twins, paired per person.
| Item | Public | Private | Gap | Same |
|---|---|---|---|---|
| Voting | 3.74 | 3.72 | +0.03 | 87% |
| Charity | 2.63 | 2.62 | +0.01 | 85% |
| Green premium | 2.77 | 2.77 | −0.00 | 89% |
| Gym | 2.61 | 2.60 | +0.01 | 84% |
Mean rating, 1–5, read from each twin’s own words.
Set this null beside the banking deep dive, and the boundary is clean: a say-do gap appears when doing the thing is expensive — not just because a topic feels sensitive.
A null result earns trust only if its limits are on the table too.
The honest claim
Brox.AI — predicting what people do, not just what they say.
~546 US digital twins, within-subject SAY vs DO, June 2026. Ratings parsed from stated text responses. Banking comparison: J.D. Power U.S. Retail Banking Satisfaction Study (Q3 2025). brox.ai · Full data & reasoning traces available on request.