03.03Fieldwork and Data CollectionAvailable

Stimulus and Concept Preparation

Prepare concepts, ads, packs and prototypes so respondents react to the idea you meant to test, not to how well it happened to be written.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Concept testing is easy to run and easy to invalidate, and the invalidation is silent. The commonest failure is parity: one concept in the set is written with more benefit specificity, more enthusiasm, a stronger reason to believe or simply more words, and it wins. The result is real, repeatable and entirely about the writing, and nothing in the data reveals it. The second commonest is fidelity mismatch: a finished, art-directed execution is used to test an idea, respondents react to the photography and the typeface, and the debrief reports that the idea was rejected.

This skill prevents both. It separates the four layers a respondent actually reacts to (idea, expression, execution, exposure context), sets fidelity to match the decision rather than the budget, imposes a common template and a six-dimension parity audit conducted by someone who did not write the concepts, chooses between monadic, sequential monadic and comparative exposure with the limitations of each stated, sets how many items one respondent can meaningfully assess, controls rotation and position, specifies the competitive frame, exposure duration and price and brand status, and pilots for comprehension before anything is fielded.

It also fixes the reporting rule that concept tests rank options reliably and predict market outcome poorly.

Best used for

  • Preparing new product or service concepts for quantitative screening
  • Preparing advertising or communications material for pre-testing
  • Preparing packaging with a realistic competitive frame
  • Preparing prototypes and interfaces for evaluation
  • Choosing between monadic, sequential monadic and comparative exposure
  • Auditing a concept set where one option is winning suspiciously

Typical inputs

What you give it.

The decision the test informs, Raw material for each item (propositions, scripts, layouts, sketches, pack designs), What is intended to vary across the set and what is held constant, Target audience and their category familiarity, Competitive set (optional), Price and brand decisions (optional), Real-world exposure conditions (optional), Norms from comparable tests (optional)

Typical outputs

What you get back.

Stimulus brief naming the layer under test and the controls on the other layers, Fidelity decision and rationale, Final versioned stimulus set in a common template, Parity audit table across length, benefit specificity, evaluative language, reason to believe, concreteness and framing, Exposure design specification (monadic, sequential monadic or comparative) with rotation or block scheme, Exposure condition specification covering context, duration, re-exposure, branding and price, Comprehension pilot results and post-change re-audit, Version log, Reporting caveats drafted before results exist

Method coverage

What the skill works through.

  1. The four layers a respondent is actually reacting to
  2. Choosing fidelity: what rough and finished stimulus each test
  3. Writing a concept set at parity, and the six-dimension audit
  4. Monadic, sequential monadic and comparative exposure, and what each supports
  5. How many concepts one respondent can meaningfully assess
  6. Rotation, position effects and balanced block designs
  7. Competitive context, exposure duration and re-exposure
  8. Price and branding: decide once, apply everywhere
  9. Comprehension piloting and why a middling score can mean confusion
  10. Why concept scores rank well and forecast badly

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

Why does the same concept keep winning every test?

Check the writing before the idea. A concept with more words, a quantified benefit, a reason to believe its rivals lack, more evaluative adjectives, or an opening that frames a problem before offering the solution will outperform, and nothing in the data will say so. Audit the set in a column, dimension by dimension, using someone who did not write it. Expect the strongest-written concept to be the one its author already preferred.

How finished should test stimulus be?

Match fidelity to the decision, and keep it uniform across the set. Rough stimulus isolates the idea, is cheap enough to test many options, and invites respondents to imagine, which produces generous and unstable ratings. Finished stimulus is closer to what will exist and confounds idea with execution, because nobody can separate a good idea rendered badly from a bad idea. A set containing one finished item among roughs is not a test of ideas; it is a test of finish.

Should I use monadic or sequential monadic exposure?

Monadic, where each respondent sees one concept, gives the cleanest absolute read, is comparable to monadic norms and approximates how people actually encounter a proposition, at the cost of a full sample per cell and no within-respondent comparison. Sequential monadic is efficient and gives within-respondent comparison, at the cost of calibration against earlier items and fatigue in later ones. Comparative side-by-side is the most sensitive for close variants and the least realistic, because it creates a choice the market does not offer.

How many concepts can one respondent evaluate?

Three to five with full diagnostics is a common working range, fewer when items are long or complex, and it should be calibrated in pilot by looking at reading time and open-end length by position rather than taken as fixed. Where more concepts must be covered, use a balanced incomplete block design so each item appears equally often and equally often in each position.

Does adding a price and competitors lower concept scores?

Usually, and substantially. That is the point. Scores fall because the test moved closer to the world, and the differentiation between concepts generally improves at the same time. The real question is never whether someone likes a concept but whether it pulls them away from what they already do, and that question needs a competitive frame and a price.

What does a middling score with vague open ends mean?

Very often comprehension failure rather than rejection. A concept that is not understood scores in the middle, because respondents rate what they think it might be. This is why a comprehension pilot, in which people say back in their own words what is being offered and what is new about it, catches problems that ratings cannot.

Can a concept test predict sales?

No, and this is worth writing into the report before results exist. A concept test removes almost everything that determines real-world performance: availability, price in context, competitive response, media weight, repeat purchase, and the fact that nobody in the market will ever read the concept. It ranks options against each other reliably. Converting a concept score into a volume or share estimate requires a calibrated forecasting design and external data.

How do I test a genuinely new idea fairly?

Recognise the handicap. Familiar propositions are easier to write, easier to understand and score better, so a novel idea is systematically disadvantaged in short-exposure testing. Options include allowing more explanation and documenting the resulting parity disparity, testing it monadically against its own reference rather than in the set, extended or two-stage exposure, or qualitative evaluation before quantitative screening. Whichever is chosen, say in the report that a low score in the standard format is weak evidence for that class of idea.

Should concepts be branded or unbranded?

There is no universally correct answer, only a decision that must be made once and applied to every item. Branded stimulus measures the idea plus the brand's permission to make the claim. Unbranded stimulus measures the idea alone and overstates what an unknown entrant could achieve. What is never acceptable is one branded item in an otherwise unbranded set.

When should stimulus be changed mid-fieldwork?

Almost never, and never quietly. Changing stimulus while a study is live splits the sample into non-comparable groups. If a defect is found in field, the decision is whether to pause, and any change is a logged amendment that partitions the data into before and after. Version every item from the start, because late copy edits from outside the research team are how a report ends up describing something that was never fielded.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out