How to test a prototype with AI

Testing a prototype used to be a three-week project. Write a screener, recruit, schedule across time zones, moderate every session yourself, then spend a weekend buried in transcripts.
With AI, it's a day. Paste a link in the morning, read quantified findings by evening.
This guide walks the full loop — setup, sessions, sample size, analysis — and links ready-to-run templates for the most common specific tests: signup flows, onboarding, forms, redesigns, and more.
What does it mean to test a prototype with AI?
The method hasn't changed. You put a clickable, unfinished version of your product in front of people who match your target users, give them realistic tasks, and watch what they actually do — where they click, where they stall, where they give up — before asking what they thought.
What AI changes is who does the work. The AI authors the study, moderates every session, and synthesizes the results. You keep the two things that were always the job: choosing the right question, and deciding what to build differently afterward.
It's still not a stakeholder review — colleagues know too much. And it's still not an A/B test — you're watching individuals reason, not measuring aggregates.
Same method. Radically different economics.
How do you set up an AI prototype test?
Paste your prototype link. That's most of it.
From a Figma prototype link or a URL, the AI authors the complete study — goals, a screener that recruits the right participants, task scenarios, and a discussion guide — in about two minutes. It arrives as an editable draft, and you treat it like one: cut a task, sharpen a probe, tighten the screener. Total launch effort runs five to fifteen minutes.
One mechanical detail for Figma: paste the /proto/ share link — the playable version participants can actually click through — not the /design/ editor link.
Your edit pass is where judgment enters. Check that tasks read as scenarios, not instructions ("you've decided to try this for your team — get set up," never "click the signup button"). Check that no task borrows the interface's own words — that's a treasure hunt with the answer printed on the map. The AI drafts to these rules; you add the context only you have, like which screen the roadmap fight is actually about.
Two prototype-readiness rules still apply, because no moderator can fix a broken artifact. Wire every element your tasks touch — a dead hotspot reads as a broken product. And put real content on the screens where comprehension matters; a pricing table full of placeholder numbers tests nothing.
Who moderates the sessions?
The AI does. All of them, at once.
Each session is a moderated voice or video interview with screen capture: the participant clicks through your prototype and thinks aloud while the AI moderator watches, follows the task flow, and probes hesitations the moment they happen. No scheduling, no calendar Tetris, no moderating eight calls yourself. Sessions start when participants are ready.
The moderator's follow-up question is the single most valuable data point in a prototype test. "The pricing felt confusing" is a complaint; "I couldn't tell if per-seat meant everyone I invite" is a fix.
That follow-up is what unmoderated testing always sacrificed for scale — a pile of recordings showing confusion, and no way to ask about it. AI moderation keeps the probe and the parallelism both. For the deeper treatment of how AI moderation compares to human sessions, see AI usability testing.
How many users do you need?
The classic answer is five to eight. That number was never a methodological principle — it was a moderation-cost artifact, the most sessions one human could personally run and synthesize.
With parallel AI moderation, 20–50 sessions in a day is practical, because forty sessions take roughly the same elapsed time as four.
Scale changes what a finding is. At eight sessions, "three people hesitated on the review screen" is a hunch. At forty, findings arrive as quantified proportions of a defined sample — how many participants stalled, backtracked, or abandoned at each step. That's the difference between a highlight reel and a number a skeptical VP will accept.
Set N by the confidence the decision needs, not by logistics. The full argument — including where the old "5 users" advice still holds — is in how many users do you really need for usability testing?
How does the analysis work?
Automatically, and better than you'd do it by hand at 2 a.m.
The analysis runs multi-pass: themes are extracted across every transcript, each theme is aligned with the specific interviews that support it, insights and recommendations are drawn from the pattern, and everything rolls up into an executive summary. It's ready about fifteen minutes after the last interview ends.
Every claim carries its evidence. Each finding cites the transcript moments behind it, so "5 of 32 participants thought the draft was locked" is one click from participants saying so on video. Colleagues argue with interpretations; almost nobody argues with the recording.
Then fix the top three problems and re-run the same tasks. When a round costs a day instead of three weeks, retesting stops being a luxury — you watch completion climb instead of arguing about whether it did. How the synthesis works under the hood is covered in analyzing user interviews with AI.
Which prototype test should you run?
The loop above generalizes; the leverage is in aiming it at the specific flow in front of you. Each guide below is a complete, ready-to-run study — screener, tasks, questions, sample size — for one common situation:
- Test a Figma prototype — the mechanics of clickable-prototype sessions, hotspots and all.
- Test your signup flow — where analytics shows the drop-off but can't explain it.
- Test your onboarding flow — the stretch between account creation and first value.
- Test form usability — fields, validation, and the trust flinches that kill conversion.
- Get feedback on design mockups — static designs, before anything is clickable.
- Test a website redesign before launch — staging-build tests that catch what the team is too close to see.
- Test a new feature idea before building it — demand and comprehension, pre-code.
- Test your value proposition — whether the promise lands, separate from whether the product works.
If your situation spans two — a redesigned signup flow, say — start from the more specific guide and borrow the screener from the other.
What should you do with your next prototype?
Test it before you build it. That was always the right answer; it just used to cost three weeks, so nobody did.
Sound methodology comes built in — screener logic, unbiased task phrasing, structured guides — so the study is well-constructed whether or not you've ever run research. The whole loop, from pasted link to cited readout, is what Sera runs end to end. For where this fits in a broader research practice, running user research in a day walks the full picture.
Or skip the reading. Paste your prototype link, spend ten minutes on the draft, and read what forty real users did with your design — tonight.
Frequently asked questions
Can AI test a prototype?
Yes — the full loop. AI authors the study from a pasted Figma prototype link or URL, moderates voice and video sessions in parallel with screen capture while participants think aloud, and synthesizes every transcript into quantified, cited findings. You keep the judgment: which question to ask and what the results mean for the design.
How long does it take to test a prototype with AI?
Launch takes five to fifteen minutes of your time — the AI drafts the full study in about two, and you edit it like a document. Because AI-moderated sessions run in parallel with no scheduling, results land in under a day, with the synthesized readout ready about fifteen minutes after the last interview ends.
How many users do you need to test a prototype with AI?
20–50 sessions in a day is practical, because AI moderation runs sessions in parallel instead of one at a time on your calendar. At that scale, findings arrive as quantified proportions — how many participants stalled, backtracked, or abandoned — rather than a handful of anecdotes from the five to eight sessions one human moderator could run.
How do you test a Figma prototype with AI?
Paste the prototype's share link — the /proto/ URL, the playable version participants can click through — not the /design/ editor link. The AI authors goals, screener, tasks, and discussion guide from it in about two minutes. Participants then click through the prototype in AI-moderated sessions with screen capture, thinking aloud as they go.
Can AI moderate a prototype testing session?
Yes. An AI moderator conducts each voice or video session live: it walks the participant through the tasks, watches their screen as they click through the prototype, and probes hesitations with follow-up questions the moment they happen. Because no human calendar is involved, every session runs in parallel, unscheduled.
How does AI analyze prototype test results?
With automated multi-pass analysis. Themes are extracted across every transcript, each theme is aligned with the specific interviews that support it, insights and recommendations are drawn from the pattern, and it all rolls into an executive summary — ready about fifteen minutes after the last interview ends, every claim cited.
Keep reading
AI Research
How many users for usability testing in the AI era?
The five-user rule was a cost artifact of human moderation. With AI interviews running in parallel, 20–50 sessions fit in a day — and the answer changes.
AI Research
AI usability testing: how it works and what it answers
AI usability testing lets you launch a test in ~10 minutes, run parallel AI-moderated sessions, and read quantified, cited findings the same day.
AI Research
How to run user research in a day
Launch a study in 5–15 minutes — paste your URL, align on goals, review the guide — and read AI-built themes and an executive summary by 4pm.
AI Research
Analyzing user interviews with AI: what works in 2026
How AI interview analysis works in 2026: transcripts to themes, quotes, and recommendations — where it beats manual coding, where it fails, the workflow.
Hear an AI-moderated interview
on your own product.
Paste a URL. Sera drafts the study, recruits participants, and runs the interviews — usually within 24 hours.
Your first 7 interviews are on us — no credit card required.