Partner Content

Artificial Intelligence and Machine Learning

August 17, 2026

5 min read

Five Questions To Ask When Choosing Your Synthetic Research Partner

Five Questions To Ask When Choosing Your Synthetic Research Partner

Learn how to evaluate synthetic research tools with 5 key questions on data generation, model validation, and human insight.

The synthetic research market has exploded over the past 12 months. For many research teams, figuring out which synthetic partner best serves their needs can feel like a daunting task.

The confusion starts with the word itself. Ask researchers what "synthetic" means and the answers splinter: conversational AI, digital twins, personas, generated responses, AI insights. Each term points to a different method, and the method behind the output determines whether the data can support the decision you're trying to make.

It’s one reason why evaluating a synthetic partner must not start with brand recognition or funding rounds, but in your provider’s ability to answer these core questions clearly.

Question 1: How Do You Generate Synthetic Data?

Ask this first. At its core, every synthetic solution uses the same approach: it takes text as input, processes it through a model's learned patterns, and predicts a human-like output. A persona generator takes a general-purpose AI model and gives it detailed instructions to answer like a certain type of person. A digital twin goes a step further, feeding the model a library of real documents or past answers so it can imitate a specific individual. Both add a layer of specificity on top of a general-purpose model that everyone else can access too. That added specificity has a ceiling, though, the model underneath still wasn't built to predict things like how someone answers a survey.

At Qualtrics we take a different path. Rather than wrapping a public LLM in prompts, we fine-tuned a foundational model with the language of research itself. Each survey is a conversation between the survey writer and the respondent, human or synthetic, that captures idiosyncrasies and contradictions that real survey takers make along the way.

Question 2: What Makes Your Model Better than a Competitor’s?

Every vendor has a "secret sauce." They should be able to articulate what differentiates them, even if they don't reveal all of it.

Our synthetic model is fine-tuned on anonymized, aggregated survey responses from millions of real survey takers, collected through the Qualtrics platform. As we scaled training, we found that breadth of questions mattered more than depth of respondents. By exposing the model to many different question types, we taught it how surveys and respondents relate, which helps it avoid the repetitive, low-variety answers that plague general-use LLMs.

That tracks with a basic data science principle: a model is most accurate when the population you want to represent looks like the data it was trained on. A trillion-parameter model trained on the open internet can do a lot of things, but reproducing how real people answer a specific survey isn't one of its core strengths. Scale and specificity, together, are what differentiate this kind of model.

Question 3: What Metrics Do You Use to Validate Model Accuracy?

Different synthetic applications — personas, conversational AI, digital twins, simulated individual-level data — require different forms of validation. Ask a synthetic vendor to see a direct comparison between the AI-generated output and the human perspective it's trying to emulate. Look qualitatively for similar breadth and depth of response and quantitatively for synthetic and human distributions placed side by side.

In our own validation, Edge Audiences achieved an average deviation of 0.07 standard deviations from human response means (Cohen's d) — a 12x improvement over general-use LLMs on the same questions. But an average can hide a bad distribution, so we also checked KL divergence, mean differences, and top-box percentages, and confirmed that relationships between answers hold up well enough to support correlation matrices and segmentation.

Qualtrics Model

Question 4: When Should You Favor Human, Synthetic, or a Hybrid Approach?

Picking a vendor is only half the job. The other half is knowing which questions to trust synthetic data with in the first place. Don’t pick the vendor that is selling you synthetic for projects where it’s not appropriate.

Synthetic performs best on enduring, attitudinal questions. It's better at "How likely are you to try Brand X?" than "How was your last visit to Brand X?" In other words, likelihood, importance, future intent, and stable attitudes track well. Very specific, recent lived experiences don't.

Synthetic can stand on its own for early exploration, comparing options, message and concept iteration, and fast directional reads, where speed and flexibility matter most. Pair it with human data for high-stakes decisions, recent lived experiences, continuity with an existing tracker or trendline (comparability over time is where human data still adds real value), and specialized audiences underrepresented in the model's training data. The rule is simple: match the method to the decision you're making. That's the discipline a good vendor should be reinforcing,

Question 5: What Ethical Frameworks Guide Your Synthetic/AI Development?

As synthetic tools take on bigger questions, governance isn't optional. A vendor should be able to explain how data enters the model, who consents to that, and what happens to individual identity along the way, not just point to a list of certifications. At minimum, that means training only on anonymized, aggregated data; keeping customers in control of their own data; and building in privacy safeguards like access controls and the ability to remove data on request.

Your vendor should speak to these governance and security questions with the same clarity they bring to model architecture.

The Point is Not Replacement

As you talk with vendors, or think about incorporating AI into your own workflows, remember that these solutions should not be about replacing the good work you are already doing. Faster and cheaper aren't the only points a vendor should be highlighting. The real value is using these tools in new ways to reach insights you couldn't get to before. Come with questions, ask for the data behind any accuracy claim, and match the method to the decision at hand. That's what earns synthetic data its place in the researcher's repertoire.

artificial intelligencesynthetic datadigital twin

Comments

Comments are moderated to ensure respect towards the author and to prevent spam or self-promotion. Your comment may be edited, rejected, or approved based on these criteria. By commenting, you accept these terms and take responsibility for your contributions.

Derrick McLean, PhD

Derrick McLean, PhD

Product Scientist, Edge COE at Qualtrics

3 articles

author bio

Disclaimer

The views, opinions, data, and methodologies expressed above are those of the contributor(s) and do not necessarily reflect or represent the official policies, positions, or beliefs of Greenbook.

About partner

Qualtrics builds technology that closes experience gaps, transforming insight into impact.

More from Derrick McLean, PhD

Why the Model Behind Your Synthetic Research Tool Matters
Artificial Intelligence and Machine Learning

Partner Content

Why the Model Behind Your Synthetic Research Tool Matters

4 min read

Learn how to evaluate synthetic research tools and build confidence in AI-generated data for better business decisions.

Testing Synthetic Data Against Academic Benchmarks: A Replication Study
Data Science

Partner Content

Testing Synthetic Data Against Academic Benchmarks: A Replication Study

7 min read

Qualtrics examines how synthetic data performs against academic benchmarks, addressing trust and validation gaps in AI-driven research.

What I’ve Learned Co-Hosting the MRII Podcast: AI Presents A Unique Opportunity for Insights to Reinvent Itself
Executive Insights

What I’ve Learned Co-Hosting the MRII Podcast: AI Presents A Unique Opportunity for Insights to Reinvent Itself

Explore why AI is an opportunity for insights teams to reinvent their role, increase business impact, and improve decision-making.

Nick Graham

Nick Graham

Founder at Vertemis

Synthetic Respondents Explained: What They Are, How They Work, and When to Trust Them
Artificial Intelligence and Machine Learning

Synthetic Respondents Explained: What They Are, How They Work, and When to Trust Them

7 min read

Synthetic respondents use AI to simulate survey participants. Learn how they work, when they're accurate, and when real respondents are still essentia...

Ashley Shedlock

Ashley Shedlock

Content Producer at Greenbook

From Automation to Augmentation: How AI Can Help Researchers Think, Feel, and Decide Better
Artificial Intelligence and Machine Learning

From Automation to Augmentation: How AI Can Help Researchers Think, Feel, and Decide Better

Explore how cognitive offloading and augmented intelligence are changing market research, creativity, empathy, and decision-making.

Manuel Garcia-Garcia

Manuel Garcia-Garcia

Global Lead Science Activation, Research and Strategy at Ipsos

How Insights Drive Disruption with Vineet Mehra, Chime’s Chief Growth and Marketing Officer
Executive Insights

How Insights Drive Disruption with Vineet Mehra, Chime’s Chief Growth and Marketing Officer

Discover what the C-Suite wants from insights teams as Mehra discusses Chime's growth, career lessons, and AI's impact on research.

Ed Keller

Ed Keller

Executive Director at Market Research Institute International (MRII)

Sign Up for
Updates

Get content that matters, written by top insights industry experts, delivered right to your inbox.

67k+ subscribers