Key takeaways
- Synthetic users are the AI participants; synthetic data is the feedback they produce, such as survey answers, ratings, interview responses and diary entries.
- On Terapage, each activity lets you build personas from up to 13 attribute categories and 50 attributes in total.
- A synthetic activity with ten AI agents typically completes in about three minutes.
- Terapage estimates synthetic users can cut early-phase research costs by 50–70% and shorten timelines by two to four weeks.
- The beginner workflow is Simulate → Test → Learn → Activate: synthetic data explores, real participants validate.
“Run it synthetic first.” More and more researchers are hearing this instruction, often before anyone has explained what synthetic data actually is, how it is created, or how far its results can be trusted.
This is where platforms like Terapage come in. Its synthetic users and data solution lets researchers create AI personas, run early tests in minutes, and then validate the strongest findings with real participants, all within one research workflow.
What Are Synthetic Data and Synthetic Users?
Synthetic data (noun): research data created by AI personas rather than collected from real people. Synthetic users are the AI participants who produce it. Also called synthetic respondents or synthetic participants.
Synthetic users are AI personas that act like your target audience. Synthetic data is what they produce: survey answers, ratings, interview responses and diary entries. In short, synthetic users are the participants, and synthetic data is their feedback.
Think of it as a flight simulator for research: realistic practice before the real study.
On Terapage, these personas are based on real-world behaviour patterns, so their answers feel realistic rather than random. The Synthetic Users & Data solution helps teams test ideas before running research with real people. For a deeper look, read about the benefits of synthetic data and how to use synthetic data responsibly.

Key Terms Every Beginner Should Know
Before you run your first synthetic study, it helps to know the language. Here are the terms you will come across most often.
Synthetic dataData created by AI that mirrors how real people behave, without coming from any real individual.Synthetic usersAI-generated participants who take part in surveys, interviews or diary tasks in place of real people.AI personasThe profiles that shape each synthetic user, such as their age, location, job, lifestyle and buying habits.Synthetic respondentsAnother name for synthetic users, most often used in surveys and quantitative research.Real (primary) participant dataAnswers that come directly from real people, gathered through participant recruitment, research panels or insight communities.FidelityHow closely synthetic data matches real-world data. The higher the fidelity, the more likely your synthetic findings will hold up with real people.Hybrid researchAn approach that uses synthetic data to explore ideas early and real participants to confirm them later, often within one mixed method research workflow.



How Is Synthetic Data Generated?
You don't need a data science background to create synthetic data. On most research platforms, it comes down to four simple steps:
- Define your audience: Choose the traits that describe the people you want to study, such as age, location, profession or buying habits.
- Build AI personas: The platform creates AI participants that match those traits and behave like real people would.
- Run your activities: The personas complete the tasks you set, from surveys and rankings to journals and interviews.
- Review the results: AI analysis summarises the responses and highlights sentiment, themes and patterns for you to explore.
On Terapage, you can build personas from a wide range of traits, including professional, financial, lifestyle and even medical attributes. Each activity lets you choose up to 13 categories and 50 attributes in total, so your personas stay focused without losing the diversity your study needs.

How to Run Your First Synthetic Study on Terapage (Step-by-Step)
Terapage organises synthetic research around a simple workflow: Simulate → Test → Learn → Activate. Here is how a beginner can move through each stage, from building AI personas to validating results with real people.
Step 1 – Simulate: Build Your Synthetic Participants
Start by deciding who your synthetic participants should be. In the Synthetic Users & Data area, pick the traits that match your audience. Each category opens into specific options, such as age, gender, skin tone or fitness level, so you can describe your audience in detail.

If your audience is niche, you can add your own category. This helps when standard fields don't quite fit, for example students with specific extracurricular interests or patients following a particular treatment path.

Step 2 – Test: Choose a Synthetic Activity
Next, choose the research activity your synthetic participants will complete. Terapage offers a full gallery of synthetic research activity types, each mirroring its real-participant equivalent:
- Journal and diary studies for longitudinal, day-in-the-life responses.
- Text activities and fill-it-out forms for open-ended and structured answers.
- Image, audio and video reviews through core activities for creative and concept feedback.
- Polls and surveys for fast quantitative reads.
- Rank-it, sort-it and matrix questions for prioritisation and comparison.
- Document review for feedback on reports, policies or case files.
- AI-moderated interviews for adaptive, conversational depth.
Once launched, synthetic activities run quickly. A typical activity with ten AI agents completes in about three minutes, and Terapage sends an email notification when results are ready. Terapage estimates that synthetic users can cut early-phase research costs by 50–70% and shorten timelines by two to four weeks.



Step 3 – Learn: Analyse Results Instantly
As soon as responses arrive, Terapage's AI-powered insights get to work. The Synthetic Insights dashboard gives a quick preview of every synthetic activity, and each activity opens into entries, AI summaries, AI analysis, manual insights and word clouds.
Beginners can lean on a few core outputs: sentiment analysis to understand how personas feel, speech maps to see who said what and when, and AI Key Moments to surface the most useful excerpts automatically. Transcripts, summaries, analysis and media can all be exported and turned into reports and analysis for stakeholders.





Step 4 – Activate: Validate With Real Participants
Synthetic findings are hypotheses, so the final step is to test the strongest ones with real people. On Terapage, you can move straight from synthetic results into real-participant research without leaving the platform.
Depending on your question, that might mean a live video interview, an online focus group with up to 100 participants on video, an AI-moderated voice interview, an AI-moderated video interview, an AI telephone interview or an always-on Long-Term Insight Community. Because the platform is fully mobile-compatible, participants can join community discussions, post updates and complete AI-moderated interviews from their phones.




To find the right people, use participant recruitment services or add participants through screening, bulk import, email invitations, shareable links or integrations such as Salesforce and Databricks, all supported by streamlined onboarding.

Reward participants for their time through incentive distribution, which lets teams send, track and reconcile eGift cards, prepaid cards and other rewards from the same workflow. For practical tips, read our guide to managing participant incentives.

Finally, bring synthetic and real findings together. The Responses & Data view shows insights from every activity in one place, and Publish Your Insights turns them into professional, shareable reports.

Types of Synthetic Data Researchers Use
Synthetic data comes in a few forms. The easiest way to tell them apart is by how much real data they contain and what kind of output they give you.
Fully Synthetic Data
Fully synthetic data is created entirely by AI personas, with no real respondents involved. It is the most common starting point in market research because it is quick, keeps privacy risks low, and lets you test ideas before recruiting anyone.
On Terapage, you schedule a synthetic AI-moderated interview and it runs on its own in the background. Each AI persona completes its interview, and you can then review the responses one by one or all together.

Partially Synthetic Data
Partially synthetic data starts with a real dataset and adds or replaces selected records with synthetic ones. Researchers use it to fill gaps, such as an underrepresented age group or region, or to protect sensitive fields while keeping the overall structure of the data intact.
Hybrid Data
Hybrid data brings synthetic and real responses together in one study. You usually start with a synthetic activity to explore your ideas, then run a similar task with real participants to see what holds up. This is where mixed method research comes into its own: synthetic data helps you narrow your options, and real people confirm what matters.
On Terapage, synthetic and real activities can live in the same study and use the same analysis and reporting tools. For example, real participants can share photos, audio or videos through a multi-task activity, while synthetic tasks on the same topic run alongside it. This makes it easy to see where AI responses and real experiences agree, and where they don't.


Qualitative vs. Quantitative Synthetic Outputs
Synthetic data can give you two kinds of answers. Qualitative outputs explain the “why”, through interview answers, journal entries or written feedback. Quantitative outputs measure the “what”, through scores, rankings, poll choices and matrix ratings.
Terapage covers both within its qualitative and quantitative research tools. Some tasks give you both at once, such as a matrix rating paired with a short explanation of why the persona chose it.



Synthetic Data vs. Real Participant Data: What's the Difference?
Synthetic data and real participant data are not competitors. They answer different questions at different stages of a study, as the comparison below shows.
Synthetic data vs. real participant data at a glance
The takeaway for beginners is simple: use synthetic data to explore and real participants to validate. As Terapage frames it, synthetic users do not replace real people; they protect your budget by narrowing your pipeline to the ideas with the highest potential.
When Should Researchers Use Synthetic Data?
Synthetic data is a powerful tool when it is used for the right job. Use this quick guide to decide.
Good fit for synthetic data:
- Screening early ideas in concept testing, before you commit budget to any of them.
- Testing a survey or discussion guide to catch confusing questions before launch.
- Exploring hard-to-reach audiences, such as clinicians, senior executives or rare patient groups.
- Handling sensitive topics, where collecting real personal data carries more privacy risk.
Not a good fit for synthetic data:
- Final go/no-go decisions that need evidence from real people.
- Precise subgroup statistics or exact response distributions.
- Fast-changing public opinion, where attitudes may have shifted since the source data was created.
Beginner checklist: Before you start, answer these five questions.
- Do I need an early estimate, or precise evidence from real people?
- Is my audience hard to reach or expensive to recruit?
- Is the topic privacy-sensitive?
- Am I testing several ideas, questions or messages at once?
- Will I validate the strongest findings with real participants afterwards?
If you answered yes to most of these, synthetic data is a good place to start.
Synthetic Data Across Industries: Beginner-Friendly Examples
Synthetic data helps wherever research is slow, costly or sensitive. Here's how different industries put it to work.
Healthcare & Medicine
Healthcare studies often involve small, hard-to-reach patient groups and strict privacy rules. With synthetic participants, teams can test patient-journey questions or symptom diaries without using any real patient data.
A good place to start is a synthetic journal activity. Once the questions are refined, the study can move to real patients through an AI-moderated voice interview, backed by Terapage's GDPR- and HIPAA-compliant security.



Consumer Intelligence
For consumer goods and services brands, synthetic data is a fast way to screen packaging, product claims and messaging before creative budget is committed. Synthetic image, video and audio reviews show which concepts resonate and which confuse, within minutes.
The winning concepts can then move into real-world research such as in-store shopping studies, do-it-at-home product tests and digital ethnography, where consumers capture real product experiences as they happen.

Financial Services
In professional and financial services, research audiences are often busy, senior and cautious about sharing personal financial details. Synthetic personas built with financial and professional attributes let teams test how different income segments might react to a new product feature, fee structure or onboarding flow.
Rank-it, sort-it and matrix activities are especially useful here, because they show how personas prioritise features and trade-offs. The strongest options can then be validated through behavioural studies and live or AI-powered telephone interviews with real customers.
Technology & Media
Technology and media teams move fast, and synthetic data helps them keep pace. Before running user experience testing with real users, teams can stress-test app onboarding flows, feature descriptions and campaign messaging with synthetic users to catch confusion early.

Research Agencies
Research agencies often need to shape proposals before a client has approved fieldwork. Synthetic data lets agencies prototype several client concepts in parallel, strengthen their research design and arrive at the pitch with early directional evidence.
Agencies can combine synthetic pre-tests with real fieldwork in one mixed method workflow, use research templates to move quickly, manage every stage through research project management and draw on Co-Pilot Research Services for recruitment, design and analysis support.
Legal Services
Legal teams often use mock jury research to see how their arguments will land before trial. Synthetic jurors can review case files, evidence and arguments through document review and image review activities, showing which points seem convincing and which raise doubts.
The strongest arguments can then be tested with real mock jurors in a live focus group, where they discuss the case together.
HR Teams
HR teams can use synthetic data to pilot employee engagement surveys before sending them to staff. Synthetic personas built with professional and psychological attributes quickly reveal unclear wording, leading questions or topics that may feel sensitive.
Once the survey is refined, real employee feedback can be gathered through AI-moderated interviews or an ongoing insight community, where employees share their views over time.
Where Synthetic Data Fits in the Wider Research Ecosystem

Ready to run your first synthetic study?
Start your 7-day free trial Request a demo