We tested whether AI personas actually disagree with you. Most don't.
A panel can pass every demographic diversity check and still fail in a way that matters more: every persona gives functionally the same answer to the one question you actually asked.
Demographic diversity and opinion diversity are not the same thing.
A panel can pass every diversity check we run — different ages, different income tiers, a genuine skeptic in the mix — and still fail in a way that matters more: every persona gives you functionally the same answer to the one question you actually asked.
It's entirely possible to build convincing, demographically distinct personas that still collapse into a single voice the moment they're asked something concrete. This is a narrower, later-stage version of the mode-collapse problem — not about whether the panel itself is built from a real distribution, but about whether its answers, once given, actually disagree.
The check, explained plainly
After a question is answered, every response gets embedded and compared pairwise using cosine similarity. We average the pairwise scores across the whole panel and flag anything above roughly 92% as functionally identical, regardless of how differently worded the answers are on the surface.
Paraphrasing doesn't fool it, because embeddings capture meaning, not exact wording — five answers that use different sentences to say the same thing still score as highly similar.
92% is a judgment call, not a law of nature — set from looking at real examples of genuinely varied versus collapsed answers, and calibrated to catch convergence without over-flagging answers that happen to share a similar structure.
Why cosine similarity over something fancier
Cosine similarity on embeddings is simple, fast, and directly measures semantic closeness — exactly the property we care about (do these answers mean the same thing) without requiring a more elaborate, harder-to-explain model.
What this check does and doesn't tell you
It tells you whether this specific panel's answers to this specific question actually varied. It doesn't tell you whether those varied answers match what real people would say — that's a separate, external claim this check was never meant to make.
Common questions about this check
Can a panel pass demographic diversity and still fail this check?
Yes — that's the entire point of running both checks separately. Demographically distinct personas can still converge on the same opinion, especially on an easy or leading question.
Why 92% specifically?
It's a calibrated judgment call, set from examining real examples of genuinely varied versus collapsed answers — not derived from a universal statistical law.
Does this check prove the panel's answers are correct?
No — it only measures whether the answers vary from each other. Whether they match real-world opinion is a separate, external question this check doesn't address.
See this check run on your own panel
Generate a panel and see its answer-similarity score for yourself.
Start a panelMore Methodology Reports
State of AI Persona Bias
A directional look at the default bias in unconstrained AI-generated personas — young, urban, agreeable — illustrated with the same real numbers already published on our methodology page, not a new live study.
How we check panel diversity against real Census data
Every panel's age distribution is compared against real U.S. adult population estimates — and we deliberately don't run the same comparison for income, because our income labels aren't real Census brackets.