04.02Data PreparationAvailable

Data Cleaning

Fixes diagnosed data problems with documented, authorised, reversible actions, and produces a cleaning log from which the original dataset can be rebuilt exactly.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Cleaning is where research data is most often altered without anyone being able to say afterwards what changed or why. The alterations are rarely careless. An analyst fixes a typo, standardises a date, merges two spellings of a brand, drops forty inattentive-looking respondents, and picks the more plausible of two contradictory answers. Each act is defensible alone. Together they leave a dataset nobody can reconstruct, and a study nobody can check.

This skill makes every change deliberate, documented and reversible. It classifies each diagnosed issue into one of five treatments (correct, recode, flag, exclude, or leave and disclose) with a stated condition licensing each. It deduplicates on evidence rather than coincidence, harmonises categories against a canonical list, standardises formats without silent conversions, and cleans open ends within limits that preserve them as evidence. Exclusion criteria are set in writing before their effect on results is visible, exclusions require a named authorisation, and a sensitivity check reports what the cleaning did to every headline figure.

The required output is the Data Cleaning Log: original issue, action, reason, impact, reversal, authorisation. The test it must pass is reconstruction in both directions.

Best used for

  • Executing an authorised cleaning specification against a validation report
  • Producing a cleaning log that satisfies an external audit
  • Removing duplicates and fraudulent records on documented evidence
  • Harmonising fragmented brand and category lists before any frequency is run
  • Resolving contradictions with a rule rather than by plausibility
  • Reporting what the cleaning decisions did to the headline figures
  • Applying identical cleaning rules across tracker waves

Typical inputs

What you give it.

Raw dataset, preserved and unmodified, with its validation fingerprint, Validation report and severity-rated issue register from 04.01, Written exclusion criteria with the date they were agreed, Researcher authorisation for exclusions, Questionnaire, routing and codebook, Previous wave cleaning log and rules (optional), Paradata and open-end text (optional), Category master list or house code frame (optional), Analysis plan (optional)

Typical outputs

What you get back.

Cleaned analysis file, with the original preserved unchanged, Data Cleaning Log with issue, action, reason, impact, reversal and authorisation, Cleaning disclosure block for the methodology section, Exclusion accounting reconciling received to analysed with no residual, Post-exclusion subgroup base table with small-base flags, Deduplication table with match evidence and retention rule, Category mapping table including deliberate non-merges, Contradiction resolution table, Sensitivity check covering every headline measure, Verification note and consolidated review points

Method coverage

What the skill works through.

  1. What data cleaning is, and where it stops
  2. The five treatments: correct, recode, flag, exclude, leave and disclose
  3. Why exclusion criteria must be written before the results are seen
  4. Excluding on quality, never on the basis of the answers
  5. Deduplication: what counts as evidence of a duplicate
  6. Harmonising categories without merging things that are merely related
  7. Standardising formats, dates and encodings without silent conversion
  8. Cleaning open ends without altering meaning
  9. Resolving contradictions with a rule, or not at all
  10. The sensitivity check: what the cleaning did to the headline figures
  11. The Data Cleaning Log and the reconstruction test
  12. Cleaning consistently across tracker waves

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

What should a data cleaning log contain?

Six things for every action: the original issue (referenced to the validation finding), the action taken in enough detail to re-execute it, the reason (the condition that licensed that treatment), the impact on the analysis, a reversal instruction, and who authorised it. The test is not tidiness but reconstruction: someone holding the original file and the log should be able to reproduce the cleaned file exactly, and someone holding the cleaned file and the log should be able to rebuild the original.

When is it acceptable to remove respondents from a survey?

When they fail quality criteria that were written down before their effect on the results was visible, on evidence, with a named person authorising the removal. Criteria describe how someone answered: speed, invariance across a grid, patterned responding, duplication, contradictory responding. They never describe what someone answered. Removing cases because of their answers, or applying a rule only to the part of the sample producing an inconvenient result, is misconduct rather than cleaning.

Is it wrong to clean data until the results look right?

Yes, unambiguously. A dataset cleaned until a finding appears is research misconduct, whatever it is called internally. The procedural protections are simple: criteria written with a date before the effect is seen, quality-based rather than answer-based exclusion, per-decision authorisation, and a published sensitivity check showing what each discretionary decision did to the headline figures whether that is convenient or not.

What is a sensitivity check in data cleaning?

A table showing every headline measure calculated on the cleaned file and on the file before the discretionary decisions (exclusions, contradiction resolutions, judgement-based harmonisation), with the difference and both bases. It goes in the deliverable, not the working file, and it covers every headline measure rather than a selection. A large movement is a disclosure obligation, not a reason to revisit the cleaning.

How do I know if two survey records are really duplicates?

Work outward from certainty. An identical response vector across the full instrument is near-conclusive. A match on demographics plus verbatim open-end text plus answer pattern is strong. A match on demographics alone is not evidence at all, because two similar people are common in any real sample. Record for every matched set which criteria were met, which record was kept, and the retention rule, applied consistently rather than chosen case by case.

Can I edit open-ended survey responses?

Only within narrow limits, and always keeping the original. Trimming whitespace, standardising encoding, correcting an unambiguous typographical error and classifying a response as blank, gibberish or off-topic in a separate variable are permitted. Rephrasing, correcting grammar, expanding abbreviations, standardising terminology, translating and summarising are not. An open end is verbatim evidence, and a cleaned verbatim that has overwritten its original has stopped being evidence.

How should I handle contradictory answers between two questions?

Resolve only where one answer is more reliable on grounds independent of what it says: a screener verified against a sample frame, a question with an explicit reference period over one without, or a filter answer the routing already acted on. Grounds that do not qualify are which answer is more common, more plausible to the analyst, or produces a cleaner result. Where no independent ground exists, flag both, keep both, exclude the case from the analyses that depend on the contradiction, and disclose.

Should I delete cases to correct an unbalanced sample?

No. Under-representation is not a data defect and deleting from the over-represented group discards real data while leaving the same problem in a smaller sample. Composition is corrected by weighting, with the weight efficiency and effective base reported. Deleting to balance also removes the evidence of the recruitment problem that caused the imbalance.

What is the difference between cleaning and transformation?

Intent. Cleaning restores the data to what it should have been: fixing an invalid code, merging two spellings of one brand, correcting a wrongly typed field. Transformation builds new structures on correct data: collapsing five age bands into three, deriving a variable, constructing a net, reversing a scale. Keeping them separate matters because they need different justifications, and because a collapse driven by the analysis plan is legitimate while the same collapse driven by how a chart looks is not.

How do I clean tracker data consistently across waves?

Write the cleaning rules as a versioned specification and record the version applied to each wave. A rule changed mid-series creates a movement that looks exactly like real change and cannot be distinguished from it without the version record. When a rule must change, change it at a declared break, apply it backwards to the previous wave where the archive allows, publish both series for one wave, and document the difference.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out