An idea that looks brilliant in the room where it was built often falls flat the moment a real customer sees it cold, without the strategy deck or shared context that made it feel obvious. Concept testing puts that moment in front of the right people before an organisation spends real money on development, production or media. This guide covers the four core concept-testing methods, the questions that actually change a decision, a seven-step process for running one well, and how WhatsApp is changing what is practical at scale.
Most weak ideas do not look weak inside the meeting that produced them. They arrive with a polished deck, a clean mock-up, and a room full of people who already understand the strategy behind them. The customer has none of that context. They see the concept cold and decide, in seconds, whether it makes sense and whether it matters to them. Concept testing brings that moment forward, while changing the idea is still cheap.
Quick answer: concept testing is research that shows an idea, such as a product, service, feature, campaign or pricing model, to its intended audience before an organisation commits to building or launching it. A good study measures comprehension, relevance, appeal, credibility, differentiation and behavioural intent, then asks open questions that explain every score. The strongest result is not always the concept with the highest average. It is the concept with the clearest evidence for the decision the team actually needs to make.
What is concept testing?
Concept testing, sometimes called product concept testing, is a market research method used to evaluate an idea before it is fully developed or launched. Researchers show a consistent version of the concept to people in the target market, then measure and explore their reaction. The goal is to estimate whether the idea has potential and work out how to strengthen it before the cost of changing direction rises.
A concept can take many forms, including:
- a new product or service
- a feature or value proposition
- packaging or brand positioning
- an advertising route
- a pricing or subscription model
- a public programme, experience or policy idea
The stimulus can be as simple as a paragraph and a sketch, or as advanced as a clickable prototype. It only needs enough detail for someone to understand the intended value. Concept testing is distinct from usability testing: concept testing asks whether an idea is desirable and relevant, while usability testing asks whether people can use the designed experience effectively. Strong teams run both, at different stages of a project.
Concept testing examples in practice
The four methods below can feel abstract until they map onto a real decision. A few illustrative scenarios show how the choice of method tends to follow the choice being made, rather than any fixed rule:
- 01A subscription pricing model. A monadic test compares three price-and-tier combinations across separate matched samples, so each price point is judged on its own rather than against a cheaper anchor shown moments earlier.
- 02A packaging redesign. A comparative test puts two or three finished packaging routes side by side and asks shoppers to choose, because the real business decision is which single design goes into production.
- 03A new app feature. A conversational, AI-moderated test walks participants through a feature description and asks adaptive follow-ups about what problem they think it solves and what would stop them from using it, surfacing objections a rating scale alone would miss.
None of these examples is a shortcut around the seven-step process later in this guide. They simply show which method tends to fit which kind of decision.
Why concept testing matters
Concept testing reduces the risk of investing in an idea that customers misunderstand, distrust, or simply do not need strongly enough. It also gives teams evidence to improve the proposition before they commit to detailed design, development, production or a media budget.
The value goes beyond rejecting weak ideas. A good test can reveal that:
- 01the customer problem is real, but the promise is unclear
- 02the benefit is attractive, but the claim lacks proof
- 03the broad market is lukewarm while one valuable segment is highly interested
- 04customers use different, more persuasive language than the brand does
- 05two average concepts contain elements that should combine into one stronger idea
The earlier these lessons arrive, the cheaper they are to act on.
What a concept testing survey should answer
A useful study, whether it runs as a full concept testing survey or a shorter conversational test, begins with a decision, not a questionnaire. "Get feedback" is too vague. Decide what the team will do differently once the results arrive.
| Decision question | What to measure | What to probe |
|---|---|---|
| Do people understand it? | Comprehension and message takeout | "How would you explain this to someone else?" |
| Does it solve something important? | Relevance and problem intensity | "When did you last face this problem?" |
| Is the promise compelling? | Appeal and perceived benefit | "What, if anything, feels valuable here?" |
| Is it believable? | Credibility and trust | "What would you need to see to believe this claim?" |
| Is it different enough? | Uniqueness and substitutability | "What does this remind you of?" |
| Could it change behaviour? | Trial, purchase, switching or adoption intent | "What might stop you from choosing it?" |
| How should it improve? | Diagnostic ratings and open feedback | "What is the one thing you would change?" |
The open probes matter more than they look. A purchase-intent score tells you where a concept landed. It rarely tells you how to move it.
Four concept-testing methods
The four main concept-testing methods are monadic testing, sequential monadic testing, comparative testing, and qualitative or conversational testing. The right choice depends on whether the team needs clean measurement, efficient screening, a forced choice, or deeper diagnosis.
Monadic testing
Each participant evaluates one concept only, which gives the cleanest read since nothing else moves the standard for comparison. It also reduces fatigue and builds reliable norms across repeated studies. Best suited to later-stage concepts and benchmark programmes, at the cost of needing a separate matched sample for every concept.
Sequential monadic testing
Each participant evaluates two or more concepts in turn, using identical measures with a rotated presentation order. It is efficient and gives a direct within-person comparison, but the first concept can set the benchmark and later ones can start to feel repetitive. Best suited to early-stage screening across a broad concept set, at the cost of order effects and participant fatigue.
Comparative testing
Participants see concepts together and choose, rank, or allocate preference between them. This suits an explicitly competitive decision, such as picking one packaging route or one headline, but it can still produce a winner even when none of the options is genuinely strong.
Qualitative or conversational testing
Participants react in their own words while a moderator, human or AI, asks follow-up questions based on what they actually said, using text, voice notes, images or video. This is where confusion, emotional resonance, hidden objections and better customer language surface, especially before a proposition is finalised. Best suited to diagnosis and multilingual markets, though themes need systematic coding if the study needs comparable results at scale.
The strongest design usually mixes methods
Traditional concept testing often splits the job in two: a large survey identifies the winner, then a handful of interviews try to explain why. That works, but it creates a weak seam, because the people supplying the score are rarely the same people explaining it.
A conversational test can collect both from the same participant, in one pass:
- 01Show the concept.
- 02Capture the immediate, unprompted reaction.
- 03Measure the core diagnostic scores.
- 04Ask adaptive follow-ups based on those scores and the participant's own words.
- 05Compare the structured data by concept and segment.
- 06Analyse the open responses to explain the differences.
This is where a WhatsApp-native, AI-moderated approach fits naturally. A study can open with structured ratings and choice questions, then move directly into an AI-moderated interview that uses each participant's own earlier answers to generate a personalised follow-up, in the same conversation as the rating. The result is quantitative comparability and qualitative explanation from the same respondent, not two disconnected phases of research.
How to run a concept test in seven steps
To run a concept test, define the decision, recruit the right audience, standardise the concepts, capture an unprompted reaction, measure a focused set of metrics, probe the reasons behind each response, and translate the evidence into a clear action.
Write the decision before the study
State the decision in one sentence: select a route, diagnose one proposition, choose between features, refine a claim, or decide whether to proceed at all. If stakeholders cannot agree on the decision, no questionnaire will rescue the project.
Recruit people who could realistically choose it
Do not test a specialist product against a convenient general-population sample. Use behavioural screeners such as recent category use, a relevant purchase, decision responsibility, or evidence of the underlying problem. Current customers are useful, but not a substitute for non-customers when growth depends on switching or category entry, so analyse the two groups separately.
Make the concepts comparable
Give every concept the same information architecture and visual fidelity. A concept board typically covers the customer problem, the core promise, how the concept works, the main reasons to believe it, the intended user, and price where price is part of the decision. If a board needs three minutes of explanation from the team, it is not yet clear enough to test.
Capture the unprompted reaction first
Ask what stood out, what the person thinks the concept is, and how they would describe it to someone else, before any diagnostic scale reveals what the research team cares about. That first response shows the hierarchy the participant actually created, not the one the team intended.
Measure a small set of decision-linked metrics
Use five to seven core measures rather than twenty near-identical scales, typically clarity, relevance, appeal, uniqueness, credibility, value for money, and likelihood to try, buy, switch, recommend or adopt. Use the same wording and scale across every concept, and keep the questionnaire short enough that the final concept gets the same attention as the first.
Probe the reason behind the score
Ask for a specific recent example rather than a general opinion. Useful adaptive probes include what made you give that score rather than a higher one, which part feels most useful and why, what feels unclear or missing, what would make you trust this promise, what do you use instead today, and what might stop you from trying it. Voice notes are particularly useful here, since tone, hesitation and emphasis often expose a tension a transcript alone flattens. On a WhatsApp-based study, a participant can rate a concept's credibility and immediately receive a tailored question about what felt doubtful, replying by text or voice note in whichever language feels natural, which keeps the explanation traceable to the score without the scheduling friction of a separate interview.
Turn the result into a decision, not a leaderboard
Build a simple decision table for every concept rather than ranking by average score alone.
| Outcome | Meaning | Action |
|---|---|---|
| Strong scores, clear explanation | The proposition travels without help | Progress and validate execution |
| Strong appeal, weak credibility | People want the outcome but doubt the promise | Strengthen the proof, mechanism or claim |
| Clear but irrelevant | People understand it; they just do not need it enough | Revisit the audience, occasion or problem |
| Relevant but confusing | The need is real; the expression is failing | Rewrite and retest |
| Polarising by segment | The average is hiding a valuable audience | Focus the proposition and the targeting |
| Weak across metrics and language | The issue is likely the idea, not the copy | Stop, combine, or return to discovery |
A "no" is a successful result when it arrives before development, media, inventory or reputation has been spent.
How many participants do you need?
There is no honest universal sample size for a concept test. It depends on whether the study is learning, estimating or comparing.
| Study goal | Practical starting point | What it supports |
|---|---|---|
| Early qualitative learning | 8–15 per priority audience | Major comprehension issues, language, unmet needs |
| Directional mixed-method test | 30–75 per concept or key cell | Patterns, optimisation, large differences |
| Quantitative comparison | 100–200+ per concept cell | More stable estimates and subgroup reads |
| Benchmarking or high-stakes launch | Calculated from required precision and expected effect | A defensible statistical comparison |
These are planning ranges, not statistical guarantees. If a one-point difference will decide a large investment, size the sample around the smallest difference that would actually change the decision, the same logic that governs qualitative interview sample sizes, where saturation, not a fixed number, decides when to stop. For multi-market studies, calculate the required base per market or priority segment rather than treating the total sample as one usable number. Three hundred participants spread across six countries does not give the same precision as three hundred participants in each.
Common concept-testing mistakes
- 01Testing the execution instead of the idea. Unequal visuals, copy length, price detail or proof points create an unfair test. Standardise what can be standardised, and be explicit with participants about what is actually being judged.
- 02Asking whether people "like" it. Likeability feels pleasant but predicts weakly. Someone can like an idea they would never choose. Anchor questions to real category behaviour instead: trial, switching, enquiry, sign-up, recommendation or purchase.
- 03Leading with the brand's explanation. A long introduction teaches participants how to interpret the concept before they have formed their own view. Let the stimulus do the work, then test what actually landed.
- 04Declaring a winner from tiny score differences. A score of 7.2 is not meaningfully stronger than 7.0 without uncertainty, sample context and a consistent qualitative explanation behind it. Precision on a dashboard is not the same thing as certainty in a decision.
- 05Averaging away the opportunity. A concept that is mediocre overall but exceptional for one commercially important segment can be worth more than the broad average winner. Read results by audience, current behaviour, need state, market and language before ranking anything.
- 06Confusing stated intent with demand. A concept test measures reaction under research conditions. It does not reproduce a shelf, a budget, a competitor's offer, inertia or real risk. Treat purchase intent as evidence, not as a sales forecast, and where it matters, follow up with a behavioural test such as a landing-page sign-up, pre-order or pilot.
- 07Stopping at the first answer. "It's convenient" is often shorthand. Convenient because it saves time? Avoids embarrassment? Feels safer? Reduces uncertainty? The first answer names the territory. The follow-up question reveals the actual driver.
Why WhatsApp changes the concept-testing workflow
Concept testing works best when the idea reaches the right people, in a format they will actually engage with, with room to explain themselves. WhatsApp removes three common sources of friction: a new app, a separate login, and a scheduled research call. Participants can view a concept, answer by text or voice note, pause, and come back when it suits them, which matters most in mobile-first and multilingual markets where email-led samples can systematically overrepresent the most connected consumers. WhatsApp's own message open rates run well ahead of email; see Yazi's own WhatsApp survey response-rate benchmarks for the full comparison.
For researchers, the bigger shift is methodological rather than logistical. The same study can collect comparable scores and moderator-style follow-ups at scale, and because the format is asynchronous, participants respond in their own time rather than on a shared call. An AI interviewer can ask why a promise feels unbelievable, what an unfamiliar word means to a participant, or which part of the idea they would keep, without forcing every respondent through the same long script.
The goal is not to replace judgement with automation. It is to spend human judgement where it matters most, designing the test and interpreting the tension in the data, while the platform handles the repetitive work of prompting, transcribing, translating and organising responses.
Yazi is a WhatsApp-native research platform built for this kind of mixed-method study. Its concept and prototype testing work runs on the same AI-moderated engine documented in Yazi's AI Interviewer: studies can run dozens of conversations in parallel, transcribe and analyse voice notes in 100+ languages, and field in as little as 24 to 72 hours once a study is configured, letting teams test, improve and retest while an idea is still flexible.
What a Yazi concept test can include
- Concept images, storyboards, packaging, claims, copy or prototype links sent through WhatsApp.
- Single-select, ranking, rating and open-ended survey questions.
- AI-generated probes informed by each participant's own answers.
- Text, voice-note, photo and video responses.
- Automatic transcription and multilingual analysis.
- Human researcher takeover when a response needs manual follow-up.
- Results compared by concept, market, audience and behavioural segment.
This is not only a faster way to collect feedback. It changes what is practical: a team can hear the reasoning of dozens or hundreds of people rather than relying on a handful of interviews to explain a much larger survey.
The practical rule
Test early enough to change the idea, and rigorously enough to trust the change. Use monadic testing when clean, independent measurement matters most. Use sequential monadic testing when early screening efficiency matters most. Use direct comparison when the real decision is a forced choice. Add conversational depth whenever the team needs to know what to fix, not merely what won.
Above all, do not ask customers to approve the work. Ask them to show how the idea fits, or fails to fit, their lives. That is the point of concept testing: not making the meeting feel safer, but making the next investment smarter.
Frequently asked questions about concept testing
What is concept testing in market research?
Concept testing is a research method that puts a proposed product, service, feature, campaign or value proposition in front of its intended audience before an organisation commits to building or launching it. It measures how well people understand and value the idea, and diagnoses what should change before development begins.
What are the main types of concept testing?
The four main types are monadic testing, where each participant evaluates one concept only; sequential monadic testing, where each participant evaluates several concepts in turn; comparative testing, where concepts are judged side by side; and qualitative or conversational testing, which explores the reasoning behind participants' reactions in their own words.
What questions should a concept test include?
A strong concept test measures comprehension, relevance, appeal, uniqueness, credibility, value and likelihood to act, alongside open questions such as how would you describe this idea, what feels most valuable, what is unclear, and what might stop you from choosing it.
What is the difference between concept testing and product testing?
Concept testing evaluates the promise and desirability of an idea before it is fully built. Product testing evaluates an actual or near-finished product, including its performance and real-world experience. Concept testing asks whether a team should build something; product testing asks how well the built version works.
What is the difference between concept testing and usability testing?
Concept testing checks whether an idea is understood, relevant and desirable. Usability testing observes whether people can complete tasks with a product or prototype. A concept can be desirable but hard to use, or easy to use but not valuable enough, which is why strong development programmes test both.
How many concepts should be tested at once?
In a sequential monadic study, three to four concepts is usually a practical ceiling before fatigue and comparison effects set in. Larger concept sets are better screened in shorter early rounds or split across matched samples. In a monadic design, each participant evaluates only one concept, so the ceiling does not apply in the same way.
Can concept testing be qualitative and quantitative?
Yes. A mixed-method concept test combines comparable ratings or choices with open-ended explanation. This tends to be more actionable than either method alone, because a team can see which concept performed best and understand why participants reacted the way they did.
Can AI be used for concept testing?
Yes. AI-moderated interviews can ask adaptive follow-up questions, transcribe voice responses, translate multilingual feedback and organise themes at scale. Researchers should still control the study design, review the evidence and make the final call on what the results mean.
How does Yazi run concept testing?
Yazi runs concept tests over WhatsApp. Participants view the concept, answer structured questions, then receive personalised AI-moderated follow-ups by text or voice note. Research teams get standardised scores for comparison, alongside the participants' own language and reasoning behind each score.
How many participants do you need for a concept test?
It depends on the goal. Early qualitative learning typically needs 8 to 15 participants per priority audience, a directional mixed-method test 30 to 75 per concept, and a quantitative comparison 100 to 200 or more per concept cell. Multi-market studies need that base calculated per market, not as one combined total.
Structured concept scores and AI-moderated follow-ups, in the same WhatsApp conversation.
Compare ideas across your own customers or Yazi's vetted panel of 1.4 million+ people across 13 African countries, and keep the reasoning behind every score.
Book a Demo →%202.png)


