Where every claim comes from
Before we ask you to trust synthetic, here is the full trail: one real study end to end, human vs synthetic theme by theme, and how we audit it.
We pre-registered a prediction in our own favor. The data said no. We published it anyway.
A pre-registered study of 2,600 synthetic respondents across eight markets, sealed by hash before we saw a single data point. It set out to test whether the psychometric layer our engine adds is what makes markets differ. It is not — and we publish that, alongside the correction we made to our own first reading when the data would not support it.
The pre-registered successor to our earlier exploratory study (n=5 per market) — built to re-test our own prior claim, with controls. Read the v1 on SSRN
Read the flagship paper (v2) on SSRN
Profile divergence within a market
An automated first reading suggested that fixing a persona's value profile shifts the way it reasons. Blind human coding did not confirm it, so we treat it as an open hypothesis, not a result.
Differences between markets
They show up, but they do not come from the psychometric layer — a control cell without it shows just as much structure. They are rendered by the base model conditioned on country and language, so we present them as hypotheses to validate against humans, not as national truths.
What we found, and what we do not claim
- Sealed by SHA-256 hash before any data. Nothing was added, removed, or reformulated after.
- The central result runs against our own product — and we published it in full.
- Our first reading suggested a stronger claim; about 79% of it was measurement noise, so we withdrew it.
- We ran the human check at our own cost: 20 blind coders, 120 items. Two behaviours cleared the bar cleanly (0.93 and 0.96) — but the coders disagree with each other on four of the six, so the gate does not pass and H2/H3 stay unconfirmed. Published, not buried.
- Nor did we see the profile change which option a persona picks: a separate A/B study showed no difference. We report it as it came out.
- Data and code with fixed seeds are published — a skeptic can re-run the entire analysis.
Provisional Phase-1 result. Market labels denote the generation context, not national truths. We have since run independent blind human coding and published it in full: it did not clear the bar we set ourselves, so "validated" still does not apply — and we would rather say that than stretch the word.
From brief, to synthetic conversation, to a traceable theme.
Before we ask you to trust synthetic research, we show you where the output comes from. This is one real study, end to end.
"Why do customers stay with their insurer even when they are unhappy?"
Internal hypothesis: Price is the main barrier.
Audience: Spanish insurance customers, mixed age and income segments.
Guide: Switching triggers, trust, coverage understanding, reactions to renewal messaging.
What stops you from switching insurer, even when you are not happy?
"Price is the brake. Complexity is the fog. The fog is worse — it stops me from even seeing the road."
Complexity blocks switching more than price.
Mapped across 217 synthetic interviews — strongest where coverage feels hardest to compare.
The brief shifted before fieldwork.
From "test price messaging" to "probe clarity, trust, and rejection of scare tactics".
What 217 Spanish insurance consumers actually said
Insurance Coverage Choice — Spain. N=217 synthetic consumers. Every interview scored; 5 rejected on quality and excluded from analysis.
We needed to understand why Spanish consumers stay with their insurer even when unhappy — and what triggers switching. The internal hypothesis: price is the main barrier. We ran 217 synthetic consumer interviews in under 30 minutes.
- Fear-based upsell was the #1 rejected pattern — 125 mentions, 0 acceptances across all 217 respondents
- The market leader dominated spontaneous recall with 530 mentions — 2.2× more than its nearest competitor
- "Complexity is the fog" — consumers could not evaluate coverage options, so they stayed put by default
The real blocker was complexity and distrust of scare tactics. The brief shifted from price messaging to transparency and simplicity — before a single real interview started.
Hypothesis: price is the main barrier.
Real blocker: complexity and distrust of scare tactics.
The fieldwork guide was rewritten before a single real interview — before any of the budget was spent on the wrong questions.
The findings that survived three independent analyses
Trust in Automated vs Human Investment Advice — United Kingdom (N=67) and United States (N=69). 1,904 analysed responses. Every finding below appeared in all three independent re-analyses of the same corpus; we publish the range across the three, not a single number.
A fintech brief asked what stops people trusting automated investment advice. We ran the same guide in two markets — and then did something we had never done before: analysed each corpus three separate times, having agreed in advance to publish only what appeared in all three.
- Transparency leads the unmet-need list in both markets: "Transparent, plain-English explanations of reasoning and fees" in the UK (8 to 13 of 67) and "Transparency and clear reasoning" in the US (7 to 10 of 69)
- Hidden fees is the top barrier in the US (9 to 12 of 69) and the second in the UK — the only barrier that leads in both markets
- "Trust was just: can I explain this to myself in plain English?" — US respondent, verbatim and traceable to the turn it was said in
- UK published 14 barriers and 9 unmet needs; US published 12 and 12. Every one present in 3 of 3 analyses.
Neither market asked for better performance. Both asked to be told why — in words they could repeat to someone else.
One analysis, one number, and no way to tell a real finding from a wording accident.
Three independent analyses. A finding is published only if it appears in all three.
Every figure on this card is a range because the study was analysed three times. Regenerating the report returns the same concept list, compared by exact string.
You don't have to "believe" in synthetic.
See where it aligns with humans. And where it doesn't.
Mirror View runs the same interview guide and analysis pipeline on humans and synthetics, theme by theme: mention rates, themes, sentiment. You audit convergence and divergence, not AI claims. Our first hybrid study is designed and pre-registered; no per-study human comparison has been published yet.
Illustrative — how human and synthetic responses line up per theme in a hybrid study.
- Same interview guide and analysis pipeline for humans and synthetics — comparable outputs.
- Quotes and evidence trace back to conversation transcripts, not free-floating AI summaries.
- Use synthetic for coverage and speed; use humans when live validation is required.
Built to be audited, not believed.
Every claim on this page maps to something you can inspect — not an AI wrapper that asks for faith.
- SHQI — 12 deterministic quality metrics scored on every interview
- Mirror View — human-vs-synthetic mode on the same guide and pipeline
- Every theme linked back to the conversation that produced it
- Full transcript export — read the raw evidence yourself
- Segment purity checked on every respondent: 180 judgments on 18 real biographies, zero errors
- Reproducible report: re-analysing a study publishes the same findings; up to 1,500 responses, only what appears in three independent analyses is published
We pressure-tested this very page with QualiSynth.
Before publishing, we ran synthetic interviews with agency researchers in the US, UK and Spain — the exact buyers this page is for.
One finding: buyers thought QualiSynth analysed their own transcripts, instead of generating respondents.
That's why this page now leads with the mechanism. Directional, fast, and honest about its limits — exactly how we'd want you to use it.