15 September 2026·9 min read·Ruly Altamirano, QualiSynth

How to test whether a synthetic panel works in healthcare

Everyone shows you the studies that worked. Here is the test design, the public benchmark, and what happened when we pointed it at ourselves.

If you are considering synthetic respondents for physician research, the question you actually need answered is not whether the transcripts read well. They do. It is whether the distribution that comes out carries information about the world.

That question has a testable answer, and most vendors never test it — because the honest test is one your product can fail. This is the design we use, and what it returned when we ran it on our own engine.

Step 1 — Benchmark against records, not surveys

Most validation compares synthetic answers to survey results. That shares the self-report problem with the thing being validated: if both exaggerate in the same direction, agreement proves nothing.

Administrative records do not have that problem. For US prescribing there is a census-grade one, free and public: the CMS Medicare Part D Prescriber Public Use File. Every prescriber, every drug, every year.

Step 2 — Pick the cut the model cannot have read

A language model has read the review literature. Ask it which specialty adopted a drug class first and it will tell you correctly, because that is in the articles. That is recall, not simulation, and it proves nothing about whether the panel can stand in for fieldwork.

So test the cut nobody has written up. We chose the geographic gradient in SGLT2 inhibitor adoption within primary care alone — state by state. That gradient is a by-product of local payer policy and panel composition. It is in the administrative record and, as far as we can tell, in no review article. There is nothing to recall.

Step 3 — Seal the design before you generate anything

This is the part that separates a test from a demonstration. We wrote the hypothesis, the metric, the threshold, the sample size rule and the classification procedure into one document, hashed it with SHA-256, and published the digest — before a single interview existed.

Three entry conditions had to pass first: reproduce an already-published figure from the raw file end to end; establish the smallest sample per state that reaches 80% power at the real measured dispersion; and run a clustering pilot to check the effective sample size was not smaller than the nominal one.

GateCriterionResult
CalibrationReproduce a published figure from the raw filePass
PowerSmallest n per state at power ≥ 0.80sd = 0.0548 → 50 per state, power 0.815
ClusteringICC below 0.02150 interviews, ICC estimate truncated at zero

The document also committed us to publishing the result whatever it turned out to be. That clause is the whole point. Without it, a pre-registration is decoration.

What happened when we ran it

2,500 synthetic prescriber interviews. Fifty states, fifty respondents each, zero failures. The rank correlation with real adoption came out at −0.039. After the attenuation correction we had declared in advance, −0.163. The sealed threshold was 0.40.

The hypothesis was refuted. Our engine did not reproduce the state-level gradient.

And the correlation is not the informative part. The dispersion is. Between-state variation in the synthetic panel was 0.0699, while pure sampling noise at that sample size already accounts for 0.0685. We simulated the null directly — four thousand replications of fifty states with no between-state signal at all — and the observed dispersion landed at the 60th percentile of that null. On the logit scale, where no ceiling can compress anything, the 73rd.

In plain terms: the engine did not generate fifty populations. It generated one population fifty times.

Step 4 — Attack your own result before anyone else does

A negative result is only worth what the objections it survives are worth. Three obvious ones could have invalidated the test rather than the engine. We measured all three.

ObjectionTestOutcome
The measure was saturatedRepeat on the logit scale, where there is no ceilingStill at the 73rd percentile of the null
You worded the question badlySecond measure, a recent count aligned to what the records countBase rate fell from 0.83 to 0.68; difference stayed flat
You compared different populationsRestrict the real denominator by prescriber activity (242,081 prescribers)Real rate rises to 0.676; synthetic gives 0.68 (post-hoc comparison)

The third one is the interesting one. The objection was correct as a diagnosis — the official denominator includes minimal-panel and part-time prescribers, and our biographies were all active physicians. But correcting for it strengthened the finding rather than dissolving it: the synthetic prescribing level turns out to be consistent with the most active quartile of the real data, while state rank order survives (Spearman 0.89 to 0.97) and the refutation is unchanged under all four denominators.

One caveat we owe you on that: the quartile was identified after seeing where the synthetic figure fell. It is an observed correspondence, not a calibration test. It disposes of the objection; it does not establish calibration.

What we can say now, and what we cannot

We can say this engine, asked to produce state-level variation within one specialty, did not produce it. We can say the declared region reaches the persona — zero of 2,500 lacked one — and that what does not arrive is the behaviour associated with it.

We cannot say that synthetic populations in general fail at geography: one engine was measured. We cannot say the engine ignores geography altogether — between countries there is measurable structure, and what this study constrains is where it comes from. And we cannot say the gradient is unattainable. What we can say is that it is not obtained by asking for it.

A secondary hypothesis about specialty ordering came out in the unexpected direction, but with 18 and 19 readable responses per group the difference is not distinguishable from chance (z = 1.52, p = 0.13) and it never had a power gate. We report it as non-conclusive, not as refuted. Calling indeterminate results negative is the same error in the other direction.

The observation we did not go looking for

A second arm asked about administrative friction. The panel named prior authorization as the leading barrier and step therapy as the most common response to payers. The biographies declared a predominantly Medicare practice, so we checked it against the contemporaneous Part D formulary files — 1.6 million rows, 456 formularies.

FrictionAs reported by the panelIn the declared programme
Prior authorizationLeading barrier0.42% of plans
Step therapyMost-cited payer response3.55% of plans

This is exploratory and it carries a limit we want stated rather than discovered: the question said "a payer requirement" and did not name Medicare, so a respondent describing a commercially insured patient was answering correctly. Part of the distance may come from the wording. The defensible reading is narrow — the friction described does not correspond to the programme the panel was told it worked in.

Why we are publishing a result that makes us look worse

Because the alternative is worse. A vendor who only publishes successes is asking you to trust a filter you cannot inspect. The pre-registration committed us to publishing whatever came out, and this is what came out.

What the test also established is the part that transfers: the design works. It caught a real limit in our own engine, cheaply, before it reached a client deck.

Being precise about what runs where, because it matters. The reproducibility requirement — three independent extractions, only what survives all three, every quote verified against the transcript — runs as standard on every study of up to 1,500 responses. The sealed pre-registration, the power gate and the public benchmark run whenever the category has a record to benchmark against. Most categories do not have one, and saying otherwise would be the kind of claim this article exists to argue against.

The reason to trust a finding is not that we produced it. It is that you can check it.

Do synthetic physician panels reproduce sub-national variation in prescribing?

Pre-registered · 2,500 synthetic prescribers · 50 states · benchmarked against CMS Medicare Part D. Working paper, not peer reviewed. Cite as a preprint. · PDF

Download the preprint

Run the same test on your category

If there is an administrative record in your market, the design above transfers directly. If there is not, the sealing, the power gate and the reproducibility requirements still do.

Describe your respondents and see what comes back

QualiSynth