Most first conversations about synthetic research arrive at this point within twenty minutes: someone asks how the results compare to human research.
I understand why. If you have spent a career commissioning surveys and focus groups, human data is the reference point. New method shows up, you hold it against the reference point. That is good instinct, and I would rather talk to a buyer who demands validation than one who doesn't.
But the question contains an assumption worth examining: that human research is a fixed point. That the human number is the true number, and any other method's job is to land on it.
We just spent several weeks testing that assumption against the published record. The result is a white paper, Asking People: The current state of human research, which examines what the academic and industry literature says about the instrument of human-based research.
Falling survey response rates, rising survey bots
Pew Research Center's telephone surveys had a 36 percent response rate in 1997. By 2018 it was 6 percent. Pew's online panel reports a 92 percent response rate per wave, and a 3 percent cumulative response rate once you count the full path from sampled population to delivered answer. Same panel, both numbers true. That 97 percent silence is where non-response bias lives: the people who never answer are not a random subset of the people you meant to ask.
The people who do answer, answer constantly. Morning Consult's own panelist research found the average online panel respondent takes nearly 11 surveys a week, and one in five takes more than 25. In a 2025 academic sample, 40 percent of participants had taken seven or more surveys in the previous 24 hours. Survey fatigue at that volume is not a side effect; it is the operating condition.
A growing share of the answers involve a machine. Thirty-four percent of participants on a major research platform report using LLMs to answer open-ended questions, and the resulting text is measurably more homogeneous and more positive than human writing. At the far end: a Dartmouth researcher built an autonomous AI respondent that passed 99.8 percent of attention checks, evaded every detection method in current use, and cost about five cents per completed interview.
Then there is plain fraud. Rep Data ran one census-balanced survey across six sample sources and flagged 29 percent of respondents. A NORC literature review puts typical industry fraud at 15 to 30 percent. The standard defences underperform: in Pew's testing, 84 percent of bogus respondents passed the attention-check trap question.
Question wording bias: which human number?
Those figures describe who answers. The deeper problem for the baseline assumption is what happens even when attentive, honest humans answer.
Pew's own question-design experiments show wording and order moving results by up to 25 percentage points on identical populations. In January 2003, 68 percent of Americans favoured military action in Iraq. Add one clause about potential casualties and 43 percent favoured the same action. Both are human numbers. Both were produced by real people answering honestly.
Sample tier moves the number too. Opt-in samples average 5.8 percentage points of error against government benchmarks; probability panels average 2.6. For adults under 30, opt-in error runs above 11 points.
So asking whether synthetic research matches human research assumes there is one human number to match. There isn't. There is a distribution of possible human numbers, and the questionnaire, the sample tier, and the state of the panel ecosystem pick one. You cannot validate against a baseline that will not hold still.
Where survey methodology still holds
This is not a case for throwing human research out.
While the instrument was never broken, it was never infallible either. It always had an error budget. The difference now is that these errors are growing.
Well-run probability-based instruments that state their error remain the strongest tier of the field, and we rely on them ourselves. When we validated our US population against the University of Michigan's consumer sentiment survey, we were validating against exactly that kind of instrument: decades of consistent methodology, published error, results that get checked against economic reality every month. That is what a usable reference looks like.
The failure mode is different. It is treating any single human study, run once, through one questionnaire, on one batch of panel sample, as ground truth. The record above says that reading deserves the same scrutiny as any other measurement.
Compared to reality
There is a better validation question, and it applies to every method equally, including ours.
Not "does method B match method A," but "how does each method perform against an outcome that actually happened?" A sales figure. An election result. A prior study with known results, re-run. Outcomes don't have response rates, wording effects, or panel fatigue. They just happened.
That is the test we ask buyers to hold us to, and the test worth asking of any research supplier, human panels included. A supplier who has never benchmarked against outcomes is asking you to take the instrument on faith. As of 2026, the published record says faith is not the appropriate posture toward any instrument, ours or anyone's.
I wrote more about the method-fit version of this argument in Synthetic Versus Human Research Is the Wrong Question. The white paper is the reference document: every number in this post, with its source, plus the rest of the record. Read it here.

