We pre-registered a prediction in our own favor. The data said no. We published it.
A pre-registered study of 2,600 synthetic respondents across eight markets, sealed by hash before we saw a single data point. It showed that fixing a persona's value profile shifts how it reasons, in the predicted direction in all eight markets, and it falsified what we would have loved to claim: that the same layer is what makes markets differ. We publish both.
Before we saw a single data point, we sealed a prediction by cryptographic hash: that the psychometric layer our engine adds is what makes eight markets answer the same dilemma differently. It was the product-favorable prediction. We registered it, along with the instrument and the full analysis plan, and committed in writing to publish the result whatever it turned out to be.
The answer was no. But the study came back with two headlines, not one. This is the writeup we committed to.
This is the pre-registered successor to our earlier exploratory study (n=5 per market). We built it to close that study's declared limitations — and to re-test, with controls, our own earlier claim about where between-market differences come from.
Read the v1 on SSRNWhat the study is
The object is not the psychology of eight countries. It is a construct-validity question for the whole field of synthetic respondents: when a synthetic population differs by country, which layer of the generation procedure produced that difference? Almost nobody isolates it, so the answer is usually a guess. We built the controls to stop guessing.
What held: the profile shifts reasoning within a market
Fix a synthetic persona's value profile to the opposite pole and the way it reasons shifts in the predicted direction, in all eight markets, as read by our sealed classifier.
In practice, a profile you set leaves a readable trace in how the respondent reasons.
The full result: two axes with different sources
The differences between markets are also reproducible (effect sizes of 0.17 to 0.38, reproducible across three independent generation seeds). But they do not come from our psychometric layer. A control cell with the layer removed shows the same between-market structure; a flat-prompt version on the same base model shows no degradation either. Two independent controls, built to catch different things, point the same way.
What emerges is a clean separation between two axes that practice tends to conflate:
| Axis | What produces it | In this study |
|---|---|---|
| Profile divergence WITHIN a market | The psychometric vector (value profile) | Shift in the predicted direction in 8/8 markets (automated classifier) |
| Differences BETWEEN markets | Country + language, as rendered by the base model | Reproducible across seeds, but the psychometric layer does NOT produce it |
The part most companies cannot publish: we corrected ourselves
Our first reading of the data suggested something stronger — that the cell without the psychometric layer had more between-market structure, an "inversion." It would have been a striking result. We did not trust it. We tested whether it was an artifact of our own coder being noisier on one cell than another, and it was: about 79% of the apparent inversion was measurement noise. So we withdrew the strong claim and kept the firm, smaller one: the layer adds no between-market structure. We caught our own over-claim before publishing, not after.
What we DO claim, and what we do NOT
In two columns, because the distinction is the product:
We do claim
- That fixing the value profile shifts how a persona reasons within a market, in the predicted direction in all eight markets.
- That the between-market structure is reproducible across seeds.
- That all of this is auditable: pre-registered by hash, with 2,600 labeled responses and fixed-seed code, all open.
We do NOT claim
- That our psychometric layer captures the culture of each country. It does not — that is the prediction the data falsified. Between-market structure is rendered by the base model; a second model, given identical country and language, did not differentiate the markets at all. Market labels denote the generation context, not a country.
- That this is "validated." The blind human coding has now run — see the update at the end. Two of the six behaviours cleared the bar we had pre-registered; the gate as a whole did not, so the word still does not apply. We set that bar ourselves, before there was any data, and we report against it either way: that is what separates an auditable study from a self-reported number.
To be exact about what the psychometric layer does, because it is easy to summarize wrong: it shifts how a persona reasons within a market (8 of 8 markets, in the predicted direction), and not the structure between countries. Two distinct axes with distinct sources; conflating them is exactly the error this study exists to correct.
Why we published it
The result has an uncomfortable face for us: the layer our product adds does not do one of the things its market value would invite you to assume. We publish it anyway because a study you can audit — pre-registered by hash, data and code open, results against our own interest included — is worth more than a validation number nobody can check. A skeptic can re-run the whole analysis and see that we did not move the goalposts. That is the whole point.
And the part that holds is on the record the same way: sealed before the data, measured against its pre-registered prediction, open to re-run. We would rather be the people who publish the study that goes against part of their own product than the people who can publish nothing anyone can check.
Update, 6 September 2026 — the human coding is done
We said blind human coding was the next phase. It has now happened, and this is what it found. We recruited 27 people on Prolific and had them read 120 of the answers blind: translated into English, stripped of any cue to their market, and never told that the texts were synthetic or that the task was called coding. Before the real items, each of them answered three questions with known answers; seven failed that calibration and were excluded from the analysis, though not from payment, and the blocks they had covered were relaunched with a replacement. After those exclusions every one of the 120 texts is left with exactly two independent readings.
Two of the six behaviours came through cleanly: agreement between our sealed classifier and the human consensus reached 0.93 and 0.96. But we had written the gate as the weakest behaviour, before we had any data, and by that rule it does not pass: 0.62 against the 0.70 we had committed to. H2 and H3 therefore remain unconfirmed, exactly as the pre-registered failure clause requires.
The interesting part is why. It is not that the classifier reads the texts badly. It is that the readers disagree with each other: on four of the six behaviours, two people reading the same answer do not agree on whether the behaviour is present. Where there is no human agreement there is no criterion to measure a machine against. The problem is our taxonomy, not our coder.
That has a consequence worth stating for the whole field, and it is the most useful thing this phase produced. Our two language-model judges agreed with each other around 90% of the time on the very constructs where two people do not reach a Krippendorff alpha of 0.5. Agreement between language models is easy to obtain and easy to mistake for validity: it measures the stability of a procedure, not the existence of the category. Only blind human coding tells the two apart, which is exactly why we paid for it.
What happens next is the ordinary work of a research programme. We rewrite the definitions of the four behaviours the readers could not apply consistently, and re-code the same 120 texts — the data is already collected and sealed, so that costs a redraft and not a new study. Separately, the human respondent arm, where real people answer the same dilemma instead of coding it, is prepared and sealed; that is the step that decides whether synthetic answers converge with human ones at all. The paper is now at v2.1 with the full Phase 2 result inside it, data and code included.
The pre-registration, the 2,600 responses and the analysis code — the full trail behind this study lives on our evidence page.
See the evidenceTry the engine — no signup
QualiSynthIf you do qualitative research, we would genuinely like to know where this would mislead you. That gap is what we are after.