The ethics of simulating a customer who doesn't exist
Generating a persona that talks like a real kind of person raises real questions about representation, stereotype, and what it means to speak 'for' a group that never consented to it.
A generated persona doesn't just answer a question. It stands in for a kind of person who never agreed to that.
When a synthetic panel generates a persona described as, say, a working parent in a specific income bracket, that persona's answer implicitly claims to represent how someone like that might think. No real person in that category consented to being represented this way, and the claim can be wrong in ways that reinforce a flattened stereotype rather than a genuine range of views.
The specific risk: flattening real diversity into a caricature
A demographic category contains enormous real variation — two people who share an age, income bracket, and job title can hold entirely different opinions. A poorly-generated persona risks collapsing that real variation into a single, stereotyped voice that claims to speak for the whole category.
This is a sharper version of the mode-collapse problem discussed elsewhere on this site: here, the cost of collapse isn't just a weaker research signal, it's a representational harm to real people who share surface traits with the persona but not its flattened opinion.
What we do about it
Every panel gets a diversity check specifically to catch personas that have collapsed into a single voice — a direct, if partial, defense against exactly this flattening.
We avoid framing panel results as "what [demographic] thinks" in our own product language, favoring "this specific generated panel's answer" instead — a smaller, more honest claim.
What we can't fully solve
The underlying language model's training data itself carries real-world biases and stereotypes that no downstream diversity check can fully correct for — a panel can be internally diverse and still reflect a skewed starting distribution.
This is a known, unresolved limitation of the category, not one we claim to have solved. We think naming it is more honest than pretending it away.
This is the ethical line we try to hold: a persona's answer is a directional signal from one generated instance, not a claim to speak authoritatively for an entire real demographic.
Common questions about persona ethics
Does thepanelist claim its personas represent real demographic groups?
No — we deliberately avoid that framing. A panel's result is described as that specific generated panel's answer, not as a claim about what a real demographic group as a whole thinks.
Can a diversity check fully prevent stereotyped personas?
No. It catches personas that have collapsed into a single voice within a panel, but it can't fully correct for biases already present in the underlying model's training data.
Why write publicly about a limitation like this instead of staying quiet about it?
Because pretending the risk doesn't exist doesn't make it go away — it just means users encounter it without warning. Naming it is part of using the tool honestly.
See how the diversity check guards against flattening
Read the methodology behind catching collapsed, stereotyped panels.
Read the methodologyMore Critical Perspectives
Synthetic panels are a gut-check, not a replacement
It would be easy to let the marketing copy drift toward 'skip user research entirely.' Here's exactly what a synthetic panel can and can't tell you, and why we keep saying so even when it costs us a cleaner pitch.
AI market research: how accurate is it really?
Accuracy isn't one number. It's a set of separate, checkable claims — and most of what gets called 'accuracy' in AI research marketing is actually a claim about something else entirely.
Why we said no to fake respondent counts
It would be easy to show '10,000 simulated respondents' on a panel result. We don't, because a panel of five distinct personas and a panel of ten thousand near-identical copies would look the same on that number.