We tried to grade our synthetic users against real people. (Learn what almost killed our experiment)
We're building our synthetic user in public, start to finish.
We tried to grade our synthetic users against real people. Turns out you can't. Learn why we ditched our test.
The plan was logical. Build a synthetic panel from the persona we made in Part 2, recruit a matching human panel through Great Question, run both through the identical survey, and chart the delta.
We didn't run it.
Building in public means showing the experiments that didn't go to plan. This was one, and it taught us more than the survey would have.
Partway in, we realised the number at the end wouldn't mean anything. You can't measure how accurate a synthetic user is, because you can't measure how accurate a human one is either.
Harri Thomas, a former Meta researcher and current Chief of Staff @ Great Question put it best when I walked him through the results we'd need:
"It's neither accurate nor inaccurate. You don't know."
"Maybe that's the finding. We thought we'd be able to generate a delta between real customers and synthetic customers. And actually, looking at it, it's too difficult to compare in any meaningful way. Therefore, we don't know."
I sat with that for a day before I agreed with it. It felt like giving up on the interesting part of this whole experiment. But he's right. The question worth answering isn't "how accurate is it." It's "what is this good for, and when should you trust it?"
So instead of a survey, I talked to UX researchers already experimenting with synthetic users.
A bit more on why I ditched the accuracy test
Before I get into the conversations with several UX research leaders, I’ll dive into why we ended up ditching the accuracy test.
- No like-for-like research available. When we audited our repo for a topic to measure both panels against, most of it was too feature-specific to work as a benchmark: this person reacting to that prototype in that sprint.
- Nothing quantifiable to measure. You could lower the bar to direction (did synthetic pick the same top option) over magnitude (by how much). Fairer. But you still can’t confirm whether it’s accurate still.
- A good score wouldn't change anything. Say the output you get from a synthetic user looks really valuable. Would you then present synthetic-user findings to your leadership team as the research? None of us would.
So where do you use synthetic users then?
Everyone I spoke to had drawn the same conclusion. Synthetic users are generally good when used early, in ideation and artifact reactions, and get less trustworthy the closer you move to a real decision.
A researcher on an AI product team put it plainly: synthetic users give you a very clean answer. That word, clean, came up in almost every conversation. Real people are messy, contradictory, sometimes wrong in a myriad of ways. Synthetic users are smooth, which helps when you're pressure-testing a rough idea, but not so much for any big strategic decisions.
A hypothesis from Caitlin Sullivan, who we interviewed in Part 1, and one the conversations backed up: synthetic tracks closest on rational, operational questions (how does this workflow work, what's missing from this plan) and drifts furthest on emotional, identity-laden ones (how does it feel to have AI doing part of your job).
One of the research leads we spoke to outlined how they’re thinking about synthetic users and where they see it working best:
- Early ideation: go wild. React to as many rough concepts as you can generate. Being wrong costs almost nothing.
- Testing real artifacts: sanity-check a script or prototype with synthetic, but start bringing real people back in.
- Strategic, high-stakes work: real participants only.

Hand a synthetic user a draft survey with a few flaws baked in and it'll often catch them, the leading question, the missing option.
When synthetic users get too glossy
Here's the twist: the more polished and convincing a synthetic user gets, the more dangerous it becomes, because people stop questioning it.
One of the research leads I spoke to built a Claude skill for two of his team's real, published personas, triggered on demand so any designer or PM can say "I want to show this to one of our customers" and get a response in that persona's voice.
The first version, he built from persona summaries, sounded generic. So they fed in real verbatims and transcripts, and the personas came alive. The linguistic tics that separate one kind of technical user from another came back, when it was fed real transcripts.
That created a new problem. His word for the result was "truthy." Glossy and agreeable, without being true. He'd walked into the uncanny valley, where the output feels real enough.
A 2026 study from MIT and Penn State puts numbers behind the worry. Distilling a user's information into a profile, which is exactly what a synthetic persona is, produced the biggest jump in "agreement sycophancy": the model's tendency to tell you you're right rather than tell you you're wrong. The more you personalize it, the more it agrees with you.
"The better they get, the harder I have to caveat it. Please use this, but remember these are not real people you're talking to." - one lead UX researcher
What makes a synthetic user worth listening to
Accuracy was off the table, so I kept pushing each person on a simpler question: what makes one synthetic user good enough for you? They all landed in the same place. Grounding it in research, and whether it can show evidence and tell you it’s not confident on a finding.
A generic model will happily play a persona. Ask ChatGPT to "be a UX researcher at a B2B SaaS company" and it will, fluently, generalizing from the entire internet so you get the average of everyone.

For a synthetic user to have any value, its answers need to trace back to evidence you can inspect. It's why building them on real transcripts instead of tidy summaries works better: the verbatims are what let it echo your customer's language back to you instead of a generalisation.
The AI-product researcher wanted the same thing as a trust signal he could see, a confidence score telling you an answer is 75% backed by your data:
"Nobody wants 75% accurate data. But at least we're transparent about it. If leadership asks, I can say there are still open questions here, but it was enough for us to move quickly."
That reframe takes the pressure off the number being "right." A 75% flag isn't a failing grade, it's the synthetic user telling you to treat the answer as directional and get the rest from humans.
This is why the skill we're shipping in Part 4 has citation and a confidence threshold built in. When it doesn't have enough to back an answer, it says so, and flags the gap instead of papering over it. A flagged gap is your next research signal.
How teams are actually building and using them
Nobody we spoke to had purchased a dedicated synthetic-user tool. Everyone had built their own or was about to, for the same reason: frontier models plus their own data get them most of the way there, and a tool outside their repo can't cite anything they trust.
Here’s 3 ways we observed UX researchers were building synthetic users:
- Persona-skill. One team built one Claude skill holding two published personas, wired to triggers so anyone can pull the right voice on demand. They started from summaries, found them generic, and reinfused with verbatims. Eventually they'll split it into separate skills as they add more personas, since the context needed per persona gets too thin.
- Live retrieval of the repo. Point a skill at the most recent interviews in a window and query against those, so the persona stays current.
- Curated persona. Hand-pick the best six to eight interviews for a segment, start with one or two personas, expand as you gather more research. Some deliberately build a negative profile from churned or unhappy customers, so the synthetic user plays skeptic rather than fan.
And what they run through them:
- Script and guide testing. The most established use. Run your interview guide past synthetic before a human sees it.
- Prototype reactions. Build a rough prototype in Figma or Claude, trigger the persona, ask how they'd use it.
- Onboarding and empathy. Point a new designer at the whole repo through the persona to get a feel for the customer before their first real call.
- Artifact critique. Feed it a PRD or draft survey and ask what's wrong.
Nobody had yet presented synthetic findings to the wider team. One researcher said the day someone shares "insights from a synthetic user" will be both a sign of adoption and a scary moment.

The shared wish: making past research live on
There’s one thing we can all agree on: we’d all like research to live on and be reusable. A UX researcher at a large productivity-software company said it best:
"We know so much about our users. It's just floating in the ether, in conversations, in Slack, in decks. If all that could inform a synthetic proto-user, I'd be over the moon to interview it.”
What research could look like, with synthetic users at the center
Today a researcher is pointed at one question, answers it, moves on, and the evidence mostly serves that one decision. But stakeholder questions fall into only about five buckets: consumer profiling and demographics, market and competitor analysis, product and concept testing, brand tracking and perception, and pricing.

Know the shape of what's coming and you can collect evidence ahead of it, building a repo deep enough that a grounded synthetic user handles the directional version of most of them on day one.
"What does research look like where the primary interaction with customers is synthetic personas? That's a really interesting paradigm for UX researchers."
- Harri Thomas, Great Question

It moves the job rather than shrinking it. When ATMs arrived, the number of bank tellers went up, because banking got cheaper and people did more of it. As research gets cheaper, teams do more of it, and the scarce skill shifts from gathering data to knowing which answers to trust. That makes a researcher's judgment more valuable, not less.
One honest counterweight, named by an AI-product researcher: something real is lost when you stop talking to people. Part of what research delivers is the clip, the moment a stakeholder hears a customer's actual voice and finally gets it. A synthetic panel still can’t do that...
Where we landed
We set out to prove how close synthetic gets to real. We ended up somewhere more useful: you can't answer that, and chasing it distracts from the question that pays off. The working position:
- Use it early, or not at all. Ideation and artifact reactions, where being wrong is cheap. Not go/no-go decisions.
- Ground it your customers transcripts. No citations means there's no traceability and therefore no trust.
- Instruct AI to flag findings as low confidence. "I don't have enough data to answer this" is the most useful thing it can say.
- It won't replace human research...yet. The clean answer is a great starting point. The messy real person still changes minds, and what shifts roadmap and business decisions.
What's coming in Part 4
Next week we bring the series home. Part 4 walks through the synthetic user skill we've been building, in detail, with the citation and confidence logic baked in.
We're revealing the whole thing, and we want testers. If you'll run it against your own repo and tell us where it breaks, we'd love to have you try it!
If you want it the day it drops, sign up to follow along.




