top of page

AI Respondents for Market Research Panels: Why the Cheapest Quote Fails

Writer: Linda Orr
Linda Orr
5 days ago
12 min read

If you are speccing out market research for a product launch right now, you are about to get three or four quotes, and one of them is going to come in at a fraction of the others. Same deliverable list. Same sample size. A week of fielding instead of five. Sometimes a tenth of the price.


Ask that vendor one question: where are the respondents coming from.

If the answer involves synthetic respondents, AI panels, digital twins, simulated consumers, or any phrasing that dances around whether an actual human answered your questions, you have found the reason it is cheap. There is no panel cost and no honoraria, because there are no people.


I have been doing this work for twenty-five years, and I have watched the industry find a new way to avoid paying respondents roughly every four years. Mall intercepts gave way to cheap online panels. Then river sampling. Then buying the tail end of someone else's omnibus. Then cutting the incentive from twenty-five dollars to five and wondering why the completes arrived at three in the morning from one IP block. Every one of those moves got defended the same way, which was that the data looked close enough and the savings were real.


This one is different only in scale. Sample and honoraria are the most visible line items on a research budget, and they are also the only line items that buy you contact with a customer. Cutting them entirely is not a discount on research. It is a decision to launch on fiction, and the published evidence on that is now large enough, recent enough, and consistent enough that nobody needs to guess.


1. What did the largest test of AI respondents in market research actually find?


A team led by Tianyi Peng and Olivier Toubia at Columbia ran the definitive version of this test and published it in Science Advances. They built digital twins of real people, then compared what the twins said against what their humans said. This was 19 preregistered studies covering 164 outcomes, with twins trained on each person's prior answers to more than 500 questions, across 1,784 real participants and topics including hiring decisions, political attitudes, privacy choices, and news consumption.


Chart showing AI digital twins correlate with the real people they represent at only 0.20 on a zero-to-one scale, with supporting statistics on synthetic sample error rates, identity collapse, and AI interview dropoff.

The twins correlated with their own humans at an average coefficient of 0.20, and were only modestly more accurate than a generic base model that had none of the person's data at all.


Sit with that second finding, because it is the entire argument compressed into one line. Handing the model five hundred real answers from your actual customer barely beat handing it nothing.


They also documented five systematic distortions and named the result a funhouse mirror: insufficient individuation, stereotyping, representation bias, ideological bias, and hyper-rationality. The authors warn that deploying these systems risks smoothing out individual voices, mispredicting ignorance and irrationality, reinforcing stereotypes, and introducing new biases including pro-technology attitudes and AI favoritism.


Translate each of those into launch terms. Insufficient individuation means your segments come back looking like each other, so you cannot tell which one to build the launch around. Stereotyping means your 58-year-old suburban buyer is a caricature assembled from training data rather than a person with a household budget. Hyper-rationality means the synthetic respondent explains its purchase decision like an economist, which is the opposite of how anybody actually buys anything.


Representation bias means the voices you most needed to hear are the ones flattened first. And pro-technology bias means that if what you are launching has a screen, an app, or an AI feature anywhere in it, the model is quietly rooting for you.


That last one should end the conversation for anyone in telehealth, SaaS, connected hardware, or consumer electronics. You are paying for an unbiased read and buying a cheerleader.


2. Why does this fail worst on the exact audience you are launching to?


Here is the part that matters most for launch research specifically.


The commercial pitch for synthetic sample is strongest exactly where real sample is hardest to get. Low-incidence B2B titles. A specific clinical population. Buyers who sit at an intersection you cannot economically field, like women over 55 who own a business in a rural market, or Spanish-speaking parents of children with a particular diagnosis. Those are the cells that blow up a research budget, and they are the cells a synthetic vendor promises to fill for free.


They are also where it collapses hardest. A recent study tested standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel, producing 21 million simulated response distributions from eight different models. In real people, subgroup opinion is roughly the additive sum of its single-identity components, and it grows 2.5 times more distinctive as identities intersect. Simulated respondents showed no such composition. A single feature explained a two-feature persona's answers better than the combination in 75 to 82 percent of subgroups, a third feature added almost nothing, and the collapse survived every prompting strategy tested. The attributes the models discarded most reliably were race and religion.


So you tell the model your target is a Black Republican small business owner, and it silently picks one of those three and answers as that. Real people do not work that way.


They get more distinctive as their identities stack, not less. If the respondents underneath your study cannot hold two identities at once, then whatever buyer personas come out the other end are decoration.


3. Is this one study, or is it the whole literature?


It is the whole literature, and it predates the current news cycle by two years.

Bisbee and colleagues published in Political Analysis in 2024 on synthetic replacements for human survey data. ChatGPT recovered the survey averages reasonably well, then failed everywhere that matters: less variation than real surveys, regression coefficients that differed significantly from the human estimates, meaningful shifts from minor changes in prompt wording, and different results from the identical prompt three months later.


Right averages sitting on top of wrong variance is the most dangerous failure mode there is, because the topline passes the sniff test in the boardroom and the segmentation underneath it is invented. You will not catch it in the debrief. You will catch it eighteen months later when the segment you built the launch around does not buy.


Verasight, a verified-panel provider, has published four white papers on this. Their fourth compared 2,000 nationally representative U.S. adults against a matched synthetic dataset built with identical demographic inputs, using questions submitted by working researchers across politics, health care, technology, education, daily life, and consumer behavior. Error was highest on questions about individual experience. Whether employers respect online degrees, personal experiences with inadequate medical care, and perceptions of customer service trends all showed mean absolute errors between 27 and 33 percent. The models performed so poorly on multi-answer questions, frequently missing options that large shares of humans selected, that those questions were excluded from the main accuracy analysis.


Multi-answer questions. As in, select every brand you considered. As in, most of a category and usage study, most of a competitive set exercise, and most of a concept screen.


Lead author G. Elliott Morris noted that synthetic samples can approximate some politically polarized attitudes while struggling with questions rooted in individual experience. Individual experience is the entire job of commercial launch research.


Nobody is paying me to find out how their category feels about tariffs.


4. What about AI-moderated interviews instead of fully synthetic data?


This is a fairer question, and the honest answer has two halves.


Verasight ran a randomized controlled trial with Outset in May 2026. All 3,160 panelists started a standard survey, then were randomized either to written open ends or to an AI-moderated video interview covering the same four topics.


The depth gain is real. The AI condition produced 4.8 times as many respondent words per assigned respondent, 128 against 27, even counting every dropout as zero. On a quality score deliberately built to reward concrete detail, specific examples, personal context, and causal explanation rather than length, the AI arm delivered 1.8 times the information per assigned respondent, with roughly 80 percent of the gap coming from follow-up probing. The median written answer to a question about family finances was two words. The AI arm pulled income source and year-over-year dollar figures out of the same question.


Now the cost. Completion fell from 99.4 percent in the written arm to 40.5 percent in the AI arm, a gap of 58.9 points, with most of the loss happening at the handoff to the external platform.


Random attrition at that scale would be expensive but survivable. This was not random. Completion was predicted by attitudes measured before randomization: enjoying typed open ends added 13 points, trusting AI companies with privacy added 9.6, using AI monthly or more added 8.7, and reporting that you try to skip open-ended questions cost 20.2 points. The surviving pool ran 7.4 points more AI-optimistic than the sample it was drawn from, and demographic weighting closed only 12 to 35 percent of that tilt.


Completers were 9.7 points more likely to be men, 9 points more likely to be Black, and less likely to be Hispanic, seniors, degree holders, or from six-figure households.

The authors put the core problem precisely, which is that using an AI interviewer to measure attitudes toward AI selects on the outcome.


And the economics are not what the deck says. One clean AI interview required assigning 2.9 respondents versus 1.01 in the written condition, and 16 percent of AI completes were fraud-flagged. You are buying roughly three times the sample to get one usable transcript, then throwing away one in six of those.


The method has a legitimate place. Verasight recommends it for depth, discovery, mechanism, and hypothesis generation, and I agree. If you need to hear how twenty customers describe a problem in their own words before you write a discussion guide, this is a good tool. If you need a number you will put in a board deck and allocate budget against, no amount of qualitative richness makes up for losing six in ten people non-randomly.


5. What does the industry's own governing body say?


In May 2026 the American Association for Public Opinion Research released its Task Force report on Responsible AI Integration in Survey Research, co-chaired by David Rothschild of Microsoft Research and Jenny Marlar of Gallup. It maps AI use across the full lifecycle from questionnaire design through interviewing, coding, analysis, and reporting, and evaluates each against validity, reliability, sensitivity, and performance, with disclosure standards attached.


The report does not ban synthetic responses. It permits clearly labeled pretesting, pilot work, and exploratory diagnostics. It places synthetic responses in the highest risk tier of AI applications in survey work, and it warns specifically that the risk rises when synthetic respondents are used to populate demographic clusters that are sparse or absent in the human data.


Pin that last sentence to the wall before your vendor call, because filling the thin cells cheaply is the precise thing these products are sold to do.


6. Is a cheap human panel any safer?


No, and this is the part that catches people who think the fix is simply to buy the cheapest real sample instead.


Sean Westwood at Dartmouth published in PNAS in late 2025 on the threat large language models pose to online survey research. He built an autonomous synthetic respondent driven by a roughly 500-word persona prompt, and it evaded state-of-the-art bot detection 99.8 percent of the time. Across the seven major national polls before the 2024 election, as few as 10 to 52 fake responses at roughly five cents each would have flipped the predicted outcome, and the agent worked when prompted in Russian, Mandarin, or Korean while producing clean English answers.


So your real choice is between verified respondents and unverified respondents, and verification is now the thing you are actually paying for. Panels that do address-based recruitment, carrier-verified phone numbers, and public-record matching cost more because they are doing work. The five-dollar complete was never a bargain and it is now a liability, because you cannot tell by looking whether a human was there.

When I spec a study, the sample source and its verification method go in the scope document by name. If a vendor will not tell you how respondents are recruited and validated, that is your answer.


7. What should you ask a research vendor before you sign?


Ask where respondents come from, how they are recruited, and how they are verified. Get it in writing in the statement of work. A serious provider answers this in one paragraph without hedging.


Ask for a validation study in your category, against real respondents, on the specific question types you plan to field. Not a general accuracy claim, not a case study, not a logo wall. A holdout comparison. If a synthetic vendor cannot produce one for a category resembling yours, you have learned what you needed to know.


Ask whether your critical questions are about broad public attitudes, where these models do relatively better, or about individual experience, purchase history, consideration sets, and multi-answer behavior, where the published error rates are large enough to reverse a decision. Launch research lives almost entirely in the second category.


Ask what the decision costs if the data is wrong. A packaging tweak is one thing. Pricing a new tier, choosing which of three segments to launch into, or committing a year of media budget is another. If being off by thirty percent costs you seven figures, the honorarium line was never the expensive part of the project.


Ask who is actually doing the analysis, and whether that person will be in the room when you present the findings. A lot of cheap research is cheap because a panel platform runs the field and a junior analyst runs the tabs, and nobody senior ever looks at whether the study answers the business question. I have written separately about why research that never touches your specific situation stays unusable, and that gap is where most research budgets die.


And ask what happens to the questions your budget cannot answer. Every study has them. The useful answer is a direct list before fielding starts. The bad answer is silence, followed by a deck that quietly implies everything got covered.


8. How do you run market research that holds up for a launch?


This is where I will be direct about how I work, because the alternative to a cheap AI panel is not simply a more expensive panel.


I have a PhD in Marketing with a specialization in statistics, and I run the research myself rather than routing it to a panel platform and a tab house. That means the design, the instrument, the fielding decisions, the analysis, and the recommendation all come from the same person, and you can ask that person why a question is worded the way it is.


Most launch engagements start before any fielding, because a surprising share of what clients want from primary research is already sitting unexamined in their own transaction records, search query reports, sales call recordings, and support tickets. That work is faster, cheaper, and made of real people by construction. I would rather spend your first two weeks there and then field a smaller, sharper study than sell you a large one that duplicates what you already own.


When we do go to field, the method follows the decision. For the launch of a 3D dental imaging product, that meant qualitative work with the clinical and channel audiences that directly shaped positioning and go-to-market. For a services company entering the US yachting market, it meant more than twenty stakeholder interviews with charter owners and marketing leaders, translated into positioning and an executable market-entry strategy rather than a summary of what people said. Those are the two shapes most launch research takes, and neither one survives contact with synthetic respondents, because both depend on specific people describing specific experiences you cannot get from a model.


Deliverables are built to be used. That means a positioning recommendation you can hand to a creative team, a segment priority you can defend to a board, and a measurement plan so you can tell afterward whether the launch worked. If the study has implications for spend allocation, it connects to marketing mix modeling and analytics rather than ending at a findings deck.


You can see the full scope of market research services, and if the project sits outside a standard shape, custom project work covers instrument design, fielding, and analysis on their own.


What drives the price is incidence, method, and sample size, in that order. A hard-to-reach B2B audience costs more than a general consumer sample because reaching those people costs more. Any vendor quoting you without asking about incidence first is quoting a template.


9. Where does AI actually belong in a research project?


Plenty of places, and I use it in nearly every study I run.


Questionnaire design and pretesting, where a synthetic pass will catch a double-barreled question or broken skip logic before you spend a dollar on field. Coding open ends, which used to eat a week of analyst time. Transcription and translation. Literature and competitive scans. First-pass thematic clustering across hundreds of transcripts. Drafting the guide. Building the analysis code. Pressure-testing a model specification.

Every one of those applies AI to data that came from people, or to a work product a human then verifies. None of them ask the model to be the customer.


That distinction is the whole thing. These systems are excellent at processing evidence and poor at manufacturing it. The moment you cross from the first to the second, you are no longer running research. You are paying an articulate system to tell you a plausible story about your market, and plausible stories are the most expensive thing a growth-stage company can buy. I have written before about why the most polished output is now the least trustworthy signal, and synthetic respondents are that same problem wearing a lab coat.


The industry has spent twenty-five years finding clever ways to spend less on talking to customers. This is the cleverest one yet, and it is still the wrong place to save. Cut the deck. Cut the tab plan. Cut the twelve-question grid nobody reads. Pay the respondents.

If you are collecting quotes for launch research and want an honest read on what each approach can actually support, book a marketing strategy call. If a cheaper method will answer your question, I will tell you that too.


Sources


Comments


Contact

Thanks for submitting!

  • alt.text.label.LinkedIn
  • upwork-logo-38004EEA61-seeklogo.com
  • entre logo

©2026 by Orr Consulting. 

Orr Consulting (orr-consulting.com) is led by Linda Orr, PhD (U.S.). Not affiliated with orrconsulting.ai or Orr Group.

bottom of page