The evidence

Where every claim comes from

Before we ask you to trust synthetic, here is the full trail: one real study end to end, human vs synthetic theme by theme, and how we audit it.

Flagship study

We pre-registered a prediction in our own favor. The data said no. We published it anyway.

A pre-registered study of 2,600 synthetic respondents across eight markets, sealed by hash before we saw a single data point. It set out to test whether the psychometric layer our engine adds is what makes markets differ. It is not — and we publish that, alongside the correction we made to our own first reading when the data would not support it.

The pre-registered successor to our earlier exploratory study (n=5 per market) — built to re-test our own prior claim, with controls. Read the v1 on SSRN

Provisional — confirmatory status awaits human validation (Phase 2)
2,600
synthetic respondents
8
markets, identical stimulus
8/8
markets where the vector shifts behavior
2
failed hypotheses, published
Two axes the field tends to conflate
What the vector governs

Profile divergence within a market

Fix a persona's value profile and its behavior changes — in the predicted direction, across all eight markets. This is what a segment is, and it is our strongest result. It is what QualiSynth is built to do.

What it does not

Differences between markets

These are real, but they do not come from the psychometric layer — a control cell without it shows just as much structure. They are rendered by the base model conditioned on country and language, so we present them as hypotheses to validate against humans, not as national truths.

Two axes: the archetype moves behavior within a market (H5, 8/8); between-market structure is the same with or without the archetype.
The honest part

What we found, and what we do not claim

  • Sealed by SHA-256 hash before any data. Nothing was added, removed, or reformulated after.
  • The central result runs against our own product — and we published it in full.
  • Our first reading suggested a stronger claim; about 79% of it was measurement noise, so we withdrew it.
  • No human judges yet: interim reliability is an independent LLM coder. Human validation is Phase 2.
  • Data and code with fixed seeds are published — a skeptic can re-run the entire analysis.

Provisional Phase-1 result. Market labels denote the generation context, not national truths. Independent human validation is the next step, and only then does "validated" apply.

Evidence trail, not black-box insight

From brief, to synthetic conversation, to a traceable theme.

Before we ask you to trust synthetic research, we show you where the output comes from. This is one real study, end to end.

Study brief

"Why do customers stay with their insurer even when they are unhappy?"

Internal hypothesis: Price is the main barrier.

Audience: Spanish insurance customers, mixed age and income segments.

Guide: Switching triggers, trust, coverage understanding, reactions to renewal messaging.

Verbatim from the study
Interviewer

What stops you from switching insurer, even when you are not happy?

Synthetic respondent

"Price is the brake. Complexity is the fog. The fog is worse — it stops me from even seeing the road."

Theme detected

Complexity blocks switching more than price.

Mapped across 217 synthetic interviews — strongest where coverage feels hardest to compare.

Decision changed

The brief shifted before fieldwork.

From "test price messaging" to "probe clarity, trust, and rejection of scare tactics".

Real study · Spain

What 217 Spanish insurance consumers actually said

Insurance Coverage Choice — Spain. N=217 synthetic consumers. SHQI 0.989 (internal quality score).

We needed to understand why Spanish consumers stay with their insurer even when unhappy — and what triggers switching. The internal hypothesis: price is the main barrier. We ran 217 synthetic consumer interviews in under 30 minutes.

What the study found:
  • Fear-based upsell was the #1 rejected pattern — 125 mentions, 0 acceptances across all 217 respondents
  • The market leader dominated spontaneous recall with 530 mentions — 2.2× more than its nearest competitor
  • "Complexity is the fog" — consumers could not evaluate coverage options, so they stayed put by default
The hypothesis was wrong — price was not the primary barrier.

The real blocker was complexity and distrust of scare tactics. The brief shifted from price messaging to transparency and simplicity — before a single real interview started.

Before

Hypothesis: price is the main barrier.

After

Real blocker: complexity and distrust of scare tactics.

Result

The fieldwork guide was rewritten before a single real interview — before any of the budget was spent on the wrong questions.

Human vs synthetic

You don't have to "believe" in synthetic.

See where it aligns with humans. And where it doesn't.

Mirror View runs the same interview guide and analysis pipeline on humans and synthetics, theme by theme — mention rates, themes, sentiment. Today the human grounding is our back-test against real survey data (World Values Survey); live per-study human comparison is rolling out. You audit convergence and divergence — not AI claims.

How Mirror View reads: a hybrid study, side by side

Illustrative — how human and synthetic responses line up per theme in a hybrid study.

Illustrative — live human validation in progress.
Efficiency (driver)
Human 67%
Synthetic 61%
High convergence
Accuracy concerns (barrier)
Human 33%
Synthetic 28%
Aligned
Learning curve (barrier)
Human 33%
Synthetic 17%
Divergence — worth probing
Why this matters
  • Same interview guide and analysis pipeline for humans and synthetics — comparable outputs.
  • Quotes and evidence trace back to conversation transcripts, not free-floating AI summaries.
  • Use synthetic for coverage and speed; use humans when live validation is required.
Why researchers trust the output

Built to be audited, not believed.

Every claim on this page maps to something you can inspect — not an AI wrapper that asks for faith.

  • SHQI — 12 deterministic quality metrics scored on every interview
  • Mirror View — human-vs-synthetic mode on the same guide and pipeline
  • Every theme linked back to the conversation that produced it
  • Full transcript export — read the raw evidence yourself
  • Population back-tested against real-world survey data (World Values Survey)
  • Honest about limits — we show you where synthetic diverges from human
We use it on ourselves

We pressure-tested this very page with QualiSynth.

Before publishing, we ran synthetic interviews with agency researchers in the US, UK and Spain — the exact buyers this page is for.

Finding

One finding: buyers thought QualiSynth analysed their own transcripts, instead of generating respondents.

Resolution

That's why this page now leads with the mechanism. Directional, fast, and honest about its limits — exactly how we'd want you to use it.

Ready to run it on your own study?