04.03Data PreparationAvailable

Missing Data Handling

Separates the five things an empty cell can mean, diagnoses whether the gaps bias the figure, and reports missingness honestly rather than filling it.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Five different things arrive in a dataset as an empty cell: a respondent routed away and never asked, a respondent who skipped, a respondent who refused, a respondent who said they did not know, and a response lost to a break-off or a fault. They have nothing in common except their appearance. The first is not missing at all, and counting those respondents as missing is a base error rather than a missing-data problem. Refusals and don't knows are substantive answers, and converting them to blanks destroys real information.

This skill resolves every gap into one of the five states, then asks the question that determines the consequence: is the non-response unrelated to anything, related to something else observed, or related to the answer itself? Only the first two are recoverable. The third produces a figure biased in a knowable direction that no method inside the dataset can fix, and the honest response is to state the direction next to the finding.

The default treatment is listwise with an honest base statement, and the skill is explicit that imputation is usually not defensible in commercial research. A don't know is never a midpoint and never a zero.

Best used for

  • Deciding whether don't know sits inside or outside the base
  • Separating routed-out respondents from genuine non-response
  • Diagnosing whether non-response is related to the answer people would have given
  • Deciding whether partial completes can be included and where
  • Assessing whether a proposed imputation is defensible
  • Reporting the complete-case base before a model silently discards a third of the sample
  • Producing a missing-data disclosure a reviewer would accept

Typical inputs

What you give it.

Cleaned dataset with its cleaning log, Missing-value code scheme distinguishing the five states, Routing and filter logic, Questionnaire as fielded, including which questions offered a don't know option, Missingness map from 04.01, Paradata including break-off point and per-question timing (optional), Analysis plan, especially planned multivariate analyses (optional), Fieldwork invitation and response counts (optional), Previous wave missing-data conventions (optional)

Typical outputs

What you get back.

Missing data disclosure block, Five-state table per variable with not-asked excluded from every missing total, Base reconciliation from total, through routing, to substantive answers, Mechanism diagnosis table with evidence and direction of likely bias, Treatment decisions in the shared issue, action, reason, impact log structure, Complete-case base and profile difference for every planned multivariate analysis, Partial complete report with break-off distribution and with-and-without figures, Unit non-response statement with real counts, Imputation record where imputation was used, with flag variable and dual reporting, Statement of variables and objectives the missingness makes unanswerable

Method coverage

What the skill works through.

  1. The five things an empty cell can mean
  2. Why not-asked is not missing, and never enters a denominator
  3. Don't know and refused as answers rather than gaps
  4. Diagnosing the missingness mechanism, informally but seriously
  5. The three mechanisms and their very different consequences
  6. Listwise treatment with an honest base statement, and when it is wrong
  7. Complete-case bases in multivariate analysis
  8. When imputation is and is not defensible
  9. Why a don't know is never a midpoint and never a zero
  10. Item non-response versus unit non-response
  11. Partial completes: include by question, not by case
  12. The disclosure a missing-data decision requires

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How should I handle missing data in a survey?

Start by working out what the gaps actually are. Separate respondents who were routed away (out of base, not missing) from those who skipped, refused, said they did not know, or were lost to a break-off. Then diagnose whether the remaining non-response looks unrelated to anything, related to an observed characteristic, or related to the answer itself. For most commercial research the correct treatment is then to analyse those who answered, state the base, state the item response rate, and disclose any likely direction of bias.

Should I impute missing values?

Usually not. Imputation produces values that are indistinguishable in the file from collected ones, requires an assumption about why the data is missing that non-probability samples rarely allow anyone to test, and its main benefit is usually cosmetic. Simple methods are actively harmful: mean imputation shrinks variance, strengthens apparent relationships, and places people at the centre of a distribution when they are most likely to sit at its edges. Where imputation is genuinely warranted, flag every imputed value, report the analysis with and without it, and disclose it wherever the figures appear.

Can I treat "don't know" as the midpoint of a scale?

No. A midpoint asserts a neutral opinion the respondent did not express, and a zero asserts an absence of the thing being measured. Both invent data. A don't know is either a reported category or it is outside the substantive base. On awareness, opinion and knowledge questions it is frequently the most important answer on the page, and excluding it silently inflates every other category at the same time.

What is the difference between item non-response and unit non-response?

Item non-response is a gap inside a participating respondent's record. Unit non-response is a person who never entered the dataset. The second cannot be diagnosed from inside the data because the evidence left, so it is described rather than corrected: invitations, starts, break-offs, completions, and achieved versus intended composition. Weighting can adjust observed composition, but it does not remove non-response bias on unmeasured characteristics, and a methodology note should not imply otherwise.

How do I know if missing data is biasing my results?

Compare responders and non-responders on every observed variable, and check whether non-response on the item relates to answers on related items. Where the topic is sensitive, or where people have a reason to conceal the answer, assume the non-response is related to the answer unless the comparison shows otherwise. That case is the one that matters, because the figure is then biased in a knowable direction and nothing inside the dataset fixes it. It has to be disclosed next to the finding.

Should I include partial completes in my analysis?

The useful question is not whether but where. A respondent who answered the first two sections has given usable data for those sections and nothing beyond. Include their answers for the questions they answered, exclude them from analyses of questions they never reached, state the resulting base variation, and report the headline figures with and without partials once, because break-off is rarely random. Where people stopped is itself a finding about the questionnaire.

Why do my survey bases change from question to question?

Two reasons, and they need to be separated. Routing means different questions were asked of different people, which is by design and correct. Non-response means people who were asked did not answer, which is a decision point. A base table showing, for each question, how many were asked, how many fell into each non-response state, and how many gave a substantive answer, resolves this in one page. A base that moves is fine; a base that moves without explanation is not.

What happens to my sample size in a regression with missing data?

It can fall much further than you expect. Ten variables at 5% missing each can leave anywhere between 50% and 95% of the sample depending on whether the same people are missing throughout. Calculate the complete-case base before running the model, compare the complete-case subset with the full sample on key characteristics, and report both. Silently proceeding is the one option that is not available.

Is listwise deletion acceptable?

It is the sensible default for most commercial research, and it is an assumption rather than a neutral act: it assumes the people who did not answer would have looked like the people who did. Say so. It is the wrong choice in three specific situations: when the resulting base falls below a reportable threshold, when the missingness is related to the answer itself, and when a multivariate analysis compounds small per-variable losses into a large and non-random complete-case subset.

How much missing data is too much?

Above roughly 10% on an analysed variable, the missingness needs an explicit disclosure next to the figure. Above roughly a third, no treatment rescues the variable: complete-case analysis describes a self-selected two-thirds and imputation invents the majority of the distribution from the minority. At that point report the item response rate, describe who answered, and say the variable cannot support the analysis. That is a finding about the instrument and belongs in the methodology note.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out