Research ethics was written to govern a bounded encounter between two humans. Synthetic research keeps none of those conditions. A pillar-by-pillar account of what changes — consent, privacy, honesty, harm — with the evidence from our own studies, including where it complicates the picture.
Every pillar of research ethics — informed consent, no harm, confidentiality, honesty — was written for a bounded encounter between two people: someone answers a researcher, in a room, once, and the data sits still afterward. Synthetic research keeps none of those three conditions.
There is no human in the room — the AI moderates. There is no bounded encounter — the interview becomes a digital twin that keeps answering questions the person never saw. And the data does not sit still — it generates. The pillars still stand. What they ask of us changes. This is that account, pillar by pillar, and we've put our own evidence next to each one — including the parts that make it harder, not easier.
Traditional research ethics was built to govern a bounded encounter between two humans. We took the humans out of the room — and found the encounter had been telling us a polite lie the whole time.
Freely agree, understanding what it is and its risks — before a bounded event.
When the event is no longer bounded, what is the person consenting to? Not simply to answer questions. They consent to an AI-moderated open interview, to a twin that will answer questions they never saw, and to that twin informing future work. The consent form was designed to describe a conversation. It now has to describe a generative artifact.
What does informed consent mean when the thing you're consenting to keeps generating new outputs after you leave?
Two things push back on the pessimistic reading. First, consent stops being a one-time gate and becomes a relationship: the respondent keeps access to their own twin — to query it, benchmark it, and (under GDPR) see, port, and erase it. A model built from you that you can actually interrogate arguably makes data rights more real than a transcript you never hear about again. Second — and this surprises people — an AI moderator can return power to the respondent:
“I can tell it if I don't like the question, or if I feel like they're leading me somewhere… Anything that would normally annoy me in a conversation with a human, I can call it out and have it change. I can't tell a human it's obviously trying to manipulate me without making it weird.”
Against Belmont's respect for persons, that is not a lesser encounter — it may be a more autonomous one. The residual is comprehension. “Freely” and “informed” assume the person pictures what they're agreeing to; a candid answer given to something that feels ephemeral, but persists, is consent we should keep working to earn — not a box we can call ticked.
Keep the data safe; let no answer be traced back to one person.
We are GDPR-compliant — and for a serious conversation, compliance is the floor, not the story. The harder point sits above the legal line: a twin is a richer re-identification surface than any single datapoint. Anonymisation protects the name, not the self, and unlike a transcript that sits still, a twin persists and generates. So we treat the twin itself as personal data — erasable, able to decay, benchmarked only in aggregate.
But privacy here does more than protect. It is the condition that makes the data honest. People tell a machine what they carefully manage in front of another person. Asked, gently, what she was proud of, one participant volunteered:
“Right now I'm trying to find a job.” “Currently, right now, I live in my car, so I'm homeless.” “I have three dogs and a cat and I take care of them.”
Material poverty, disclosed without prompting, because nothing in the room was judging her and nothing could trace it socially back to her. The candour that makes synthetic research good is bought with anonymity. Which is exactly why anonymity can't be treated as a compliance checkbox — it is load-bearing for the whole method, and for the person.
Report true facts — without faking, fabricating, or dressing one thing as another.
Here is the pillar that synthetic research forces us to rethink, not just uphold. Traditional honesty means reporting what respondents said. But faithfully reporting what people say can mean faithfully reporting a number that does not exist. In our banking calibration — the same 235 people, their answers scored against real behaviour (J.D. Power, Q3 2025) — stated intent missed reality by 38.8 points. The same people's twins, reasoned through the friction of actually switching, missed by 2.5.
Banking is our reality-validated case. Behaviour-framing (DO) beat the same people surveyed by 15.5× and the best engineered LLM prompt by 12.8× — same twins, same humans, only the framing changed.
Two of those numbers are uncomfortable and both are true. Filtering the survey to its cleanest, most attentive third made the prediction worse (42.9pp) — the careful respondent, the one traditional hygiene prizes, answers the stated-intent question most earnestly, and so most wrongly. And where a twin disagreed with the very person it was built from, the twin was right 92% of the time. The authentic human voice is not automatic ground truth. So honesty, for us, means reporting the stated number, the behavioural number, and the gap between them — never the flattering single figure.
It also means never letting inference float free of evidence. Every twin response carries two fields we keep distinct: the reasoning, which is generated, and the supporting quotes that back it — lifted verbatim from the person's own interview, never rewritten. Here is one row:
Every generated inference ships with its receipts. The reasoning on the left is backed by supporting quotes pulled verbatim from the person's own interview — held in a separate field, never merged into the inference. It's a claim you can audit against real voice.
A twin can also fabricate fluently past its evidence. Honesty means calibration: a method that is directionally accurate in aggregate does not license an individual verdict. And presenting synthetic as real is the field's coming category of misconduct — one that has no shared norm yet. We think provenance labelling should be a standard, not a courtesy. We already publish it that way, losses included.
Try hard to prevent physical, mental, or social damage.
In traditional qualitative research the moderator is the safety net — they notice distress, they can stop, they can care, they can refer. This is the one pillar synthetic research genuinely weakens, and we won't pretend otherwise. The disclosures above are real people. A real person said she lives in her car. Another described putting off a colonoscopy because of the deductible. They said it to a machine, in an open interview.
Was anyone in the room? At the scale synthetic research enables — many interviews at once — the honest answer is no. Duty of care with no human witness has no traditional analog and no easy fix. It is the open work.
Two things sit beside it, fairly. Synthetic moderation also removes harms — no judgment, no awkwardness, none of the quiet coercion of being led by someone with more power in the room. And harm can travel through the output, not just the encounter: our own work shows the say/do gap is a subsidy to incumbents. Across a panel of 1,053 digital twins, 86% said they'd switch health provider for a better one; 47% actually did — yet where switching costs nothing, in retail, the gap nearly vanished (97% said, 92.5% did). Friction is the incumbent's free retention. The friction map that lets a challenger earn a switch by removing a form is the same map an incumbent could use to exploit depletion by adding one. Pointed the right way, that map does something genuinely good for people: it lowers the wall so they can finally act on what already serves them — the better plan, the fairer account — instead of paying an inertia tax to whoever they're already with. Who holds the levers is an ethics question, not a technical one.
The arc is short. The twin can be more honest than you — because you were more honest with it than you'd be with a person — and holding that gift responsibly is work we have not finished.
Traditional research ethics assumed a self-reporting human in a bounded room was the best truth available. Our own evidence says that human was off by nearly forty points, and the most earnest ones were the most wrong. That doesn't retire the pillars — consent, privacy, honesty, no harm are as binding as ever. It changes what standing behind them requires now that the room is empty: more transparency, not less; provenance kept visible; and a plain admission that on duty of care, the field — us included — still has a room to build.