- Home
- Research skills
- Sampling Strategy
Sampling Strategy
Defines the population and the frame gap, states what your sample licenses you to claim, and sizes the study from the subgroups rather than the total.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
Sampling failures are the least visible and least recoverable errors in research. A study can be well designed, well written and well analysed, and still be about the wrong people, because the population was never defined, the frame was accepted as given, or the size was set by the budget and then reported as though it had been derived.
Two failures dominate. The first is the claim mismatch: a non-probability sample reported in the language of population estimation, complete with a margin of error that is arithmetically calculable and substantively meaningless. The second, and the more common in practice, is the sizing failure where the total works and the subgroups do not. A study of 800 that everyone is happy with, until the analysis needs six sectors crossed with three size bands and every cell is unreportable.
This skill works in the right order. It defines the population operationally, measures the gap between that population and the frame you will actually use, states the licensed claim in one sentence you can put at the top of the report, and derives the total sample from the reporting cells rather than checking the cells against a total someone already chose. It covers stratification, quotas and boosts and what each obliges the analysis to do, both sizing logics, weighting and effective bases, incidence and feasibility, and who is systematically missing. For qualitative work it is direct: adequacy is governed by scope and saturation, not by power, and the proposal number is a planning estimate.
Best used for
- Sizing a study from the subgroups it must report on
- Testing whether a briefed sample size supports the analysis requested
- Stating what a non-probability sample can and cannot claim
- Defining a population and measuring the frame gap
- Designing quotas, stratification and boosts with their consequences
- Giving a defensible qualitative sample range without pretending saturation is forecastable
- Deciding weighting before fieldwork rather than after
- Establishing what an achieved sample still supports
Typical inputs
What you give it.
Population of interest with inclusion and exclusion rules, Every subgroup that will be reported or analysed separately, The frame or access route that will actually be used, The claim that must be supportable at the end, Known or estimated incidence with its source (optional), Prior variance or prior proportions on key measures (optional), The smallest difference that would change a decision (optional), Population totals for weighting, with source and date (optional), Historical response or cooperation rates (optional)
Typical outputs
What you get back.
Sampling Strategy document, Operational population definition with analysis and response units, Frame and coverage statement with the direction of likely bias, One-sentence licensed claim, Selection design, with stratification, quotas and boosts and their analysis obligations, Cell-level sizing table with required and planned bases, Sizing inputs labelled supplied, historical or assumed, Qualitative variation matrix, planned range and saturation review point, Weighting decision with effective base implications, Feasibility, incidence status and a pre-decided shortfall response, Non-response reporting plan
Method coverage
What the skill works through.
- Defining the population so two people would agree
- The frame, and the gap you cannot fix later
- Undercoverage, overcoverage, duplication and differential
- Probability and non-probability: what each licenses you to claim
- Stratification, quotas and boosts, and their analysis consequences
- Two sizing logics: precision and comparison
- The real sizing task is the subgroups
- Deriving the total from the cells
- What to do when the derived total is unaffordable
- Qualitative sample size: scope and saturation, not power
- Weighting, effective bases and what weighting does not fix
- Incidence, feasibility and deciding the shortfall response in advance
- Non-response and who is systematically missing
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
How do I calculate the sample size I need?
It depends on which logic applies. For estimation you need the precision you require, the expected value of the estimate, the confidence level and any design effect. For detecting a difference you need the smallest difference that would change a decision, the split between groups, the variability, the threshold, the power and the number of comparisons planned. Without those inputs no number is derivable, and one produced anyway is a guess with a decimal point.
Why does my sample work for the total but not for the subgroups?
Because it was sized from the total. Subgroup requirements should drive the total, not the other way round, and crossed cuts multiply fast: six sectors by three size bands is eighteen cells, and the larger groups take most of the sample. List every cut that will appear in the report, work out the base each needs, and build upward.
What can I claim from a non-probability sample?
Description of the achieved sample, and comparisons within it with the selection mechanism stated. Not a margin of error, not a confidence interval interpreted as population error, and not a population estimate presented as such. No sample size changes this, and weighting to population totals corrects composition on the weighting variables without converting the design into a probability one.
Is a quota sample representative?
It is composition-matched, which is worth having and is not the same thing. A quota fills each cell with whoever in that cell was willing and available, so it controls who is in the sample without correcting why they answered. Its particular risk is that a total which matches population figures hides non-response bias inside it.
How many interviews do I need for qualitative research?
Enough to cover the variation the question spans, and then enough more to see whether new sessions are still producing new material. Build a variation matrix of the characteristics across which experience plausibly differs, allow enough sessions per cell to see repetition, and express the answer as a range with a minimum and a defined review point. Any number fixed in advance is a planning estimate, and should be labelled as one.
What is coverage error?
The gap between the population you want to describe and the frame you can actually reach. It matters not because someone is missing but because the missing group differs on what you are measuring. Write down who is undercovered, who is overcovered, whether anyone can appear twice, and the likely direction of the resulting bias. No sample size and no weighting repairs it.
Does weighting fix a bad sample?
It corrects known composition on the variables you weight by, and it costs effective base size everywhere. It does not correct unknown selection, it does not reach people the frame never contained, and it does not turn a non-probability sample into a probability one. Report the effective base, not the respondent count, in every base and every test.
What do I do if a quota cell will not fill?
Decide before fieldwork, not in week three. The options are: extend fieldwork, which shifts the sample toward late responders; relax the quota, which changes the population that cell describes; substitute an adjacent group, which changes what the cell can be called; or report short with the base shown and the limitation attached. Each costs something. Choosing calmly in advance produces a better choice.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
