04.01Data PreparationAvailable

Data Validation

Finds every defect in a raw research dataset and rates it by severity, with counts, examples and analysis impact, without changing a single value.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Every dataset arrives looking finished. The column count is plausible, the labels read sensibly, and the first twenty rows look normal. Then the defects surface one at a time over the following fortnight: a filtered question answered by people who should never have seen it, a battery where sixty respondents chose the same point on every item, an age field containing a 3 and a 214, two records with identical open ends and different IDs, a brand list where one brand appears under six spellings, and a headline that moves five points once any of it is addressed.

This skill runs the full diagnosis in one pass and produces one document. It checks structure against the questionnaire, reconciles every base against the routing, tests values against permitted ranges, builds contradiction checks from the instrument, separates suspicious distributions from merely unusual ones, assesses response behaviour across several independent flags rather than one, maps missingness in its four distinct states, and finds fragmented category lists.

It changes nothing. Diagnosis and treatment are deliberately separated: the researcher sees the full picture, then authorises the cleaning, and the resulting log can be audited against this report item by item.

Best used for

  • Quality checking a raw dataset before any analysis begins
  • Reconciling every question base against the routing specification
  • Identifying speeders, straight-liners and duplicate respondents defensibly
  • Auditing a dataset supplied by a third party with no documentation
  • Clearing a tracker wave movement before reporting it as change
  • Building a documented quality position for a supplier or client dispute
  • Producing the diagnosis that authorises a cleaning specification

Typical inputs

What you give it.

Raw respondent-level dataset in the state it arrived, Questionnaire as fielded, with wording, response lists and scale labels, Routing, filter and skip logic, Codebook or data dictionary with valid ranges and missing-value codes, Sample and quota plan with intended base sizes, Paradata such as completion times, device type and break-offs (optional), Fieldwork log and in-field quality actions already taken (optional), Previous wave dataset and validation report (optional), Analysis plan (optional)

Typical outputs

What you get back.

File fingerprint proving the data was not altered, Validation scope statement naming checks run and checks impossible, Summary judgement on whether the data supports the planned analysis, Severity-rated issue register with counts, examples and analysis impact, Base reconciliation table against the routing, Response behaviour flag table with thresholds and flag overlap, Missingness map with not-asked, not-answered, don't know and lost separated, Category harmonisation candidates, Consolidated review points and decision requests, Statement of what could not be checked

Method coverage

What the skill works through.

  1. What data validation is, and why it changes nothing
  2. Fingerprinting the file before you touch it
  3. Structural checks: rows, columns, IDs and the question that is missing
  4. Reconciling every base against the routing
  5. Out-of-range, impossible and wrongly typed values
  6. Logic checks and contradictions between questions
  7. Suspicious distributions versus merely unusual ones
  8. Speeders, straight-liners and pattern responders
  9. Duplicate and near-duplicate respondents
  10. Mapping missingness in its four separate states
  11. Inconsistent category labels and fragmented brand lists
  12. Rating severity by analysis impact rather than by count
  13. What a validation report contains, and what it must never contain

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

What is the difference between data validation and data cleaning?

Validation finds and reports problems. Cleaning fixes them. They are separated on purpose: if the same pass diagnoses and treats, nobody can later see what the data looked like before, and nobody can demonstrate that the exclusion rule was set before its effect on the results was visible. Validation produces the diagnosis, the researcher authorises the treatment, and the cleaning log can then be audited against the validation report line by line.

How do I check whether a survey base is correct?

Derive who should have seen the question from the routing rule, then compare that set with the set that has a valid answer. Counting the valid answers and calling that the base is the commonest error in survey analysis, because it silently absorbs routing violations. Three different discrepancies matter: people who answered and should not have (a contaminated base), people who should have answered and did not (item non-response or break-off), and a base that is right in total but wrong in composition.

How do you identify poor-quality survey respondents?

Never on one flag. Compute completion speed relative to the distribution, straight-lining within each grid, patterned responding, poor or duplicated open ends, and near-duplicate respondents, then report how many respondents trigger one flag, two flags, three or more. Fast people exist and consistent people exist. The signal is convergence across independent flags. And a respondent is never flagged for what they answered, only for how they answered.

Should I remove speeders from my survey data?

That is a decision for the researcher, made against a threshold that was set before its effect on the results was seen. A validation report should give the counts at several thresholds so the choice is visible as a choice, and should note whether any quality criteria were agreed in advance. A threshold picked after the topline is known has stopped being a quality criterion.

What is straight-lining and how do I detect it?

Straight-lining is choosing the same scale point for every item in a grid. Detect it as zero or near-zero variance across the items in each battery, assessed battery by battery rather than across the whole questionnaire. It is more diagnostic in batteries containing reverse-worded items, because a genuine respondent should break the pattern there.

How do I tell a data defect from a real finding in a distribution?

Cross-check against collection metadata. Defects cluster by collection date, sample source, device or interviewer; genuine findings usually do not. Before calling a distribution suspicious, state the reason for suspicion, run the cross-check, and consider whether the instrument explains it. A spike at a mid-point is often a question with no "don't know" option, which is a finding about the questionnaire and not something to be cleaned away.

How should missing data be counted in a validation report?

As four separate things that are never summed. Not asked means routed away and is not missing at all: those people are out of base. Not answered is item non-response. Don't know and refused are substantive answers that happen to be non-substantive responses. Lost means a technical failure or break-off. A single combined percentage is worse than no figure, because it hides which of the four you have.

Why do the same brand or category names appear more than once in my data?

Differences of case, spacing, punctuation, abbreviation, spelling and language variant each create a separate category for one real thing, and the effect is a systematic undercount of whatever is fragmented. Sort category lists by normalised text rather than by frequency and inspect adjacent entries: a brand that ranks seventh under its largest variant may rank second once the variants are combined.

Can I validate a dataset without the questionnaire?

Only partially, and the report has to say so. Structural and internal-consistency checks are still possible. Routing violations, out-of-base answers, invalid codes and wrong scale directions are not, and those are where the expensive errors live. A validation is only as strong as its scope statement, and a check that could not be run must never be presented as a check that passed.

How serious is a data quality issue, and how do I decide?

By what would go wrong in the analysis if it were left untreated, not by how many cases it affects. An error touching three respondents on the primary outcome measure outranks one touching three hundred on a variable nobody will report. Rate each issue Critical, Major, Minor or Note, and write the consequence for the analysis into the rating so a reviewer can disagree with the judgement rather than merely accept it.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out