03.04Fieldwork and Data CollectionAvailable

Fieldwork Monitoring and Response Quality

Diagnose a study while it is still in field, when a defect can still be fixed rather than merely disclosed in the limitations.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Most data quality problems are found after fieldwork closes, when every option is bad. A question half the sample misread can be excluded, not fixed. A quota that filled with the wrong composition can be weighted, badly. Fieldwork is the only window in which a study can still be repaired, and it is routinely treated as a waiting period between launch and data delivery.

This skill turns that window into a set of decisions. It sets thresholds before launch so they cannot be chosen after the results are known, reads a soft launch as seven specific questions rather than a look at the data, diagnoses drop-off by pattern (a spike at a question means something different from a steady gradient), compares length in practice to length in design, tracks quota fill rates and the time confounding that comes from cells filling at different speeds, sets defensible speeder thresholds, corroborates straight-lining, classifies open-end quality including duplicate and machine-generated text, checks within-respondent contradictions, and separates fraud signals from inattention.

Its central discipline is diagnostic: before attributing any signal to respondents, cross it by question, by source and by device. Most "bad respondents" turn out to be a broken question on small screens. Every intervention is logged, because changing an instrument mid-field splits the sample into non-comparable groups.

Best used for

  • Reading a soft launch and making the release decision
  • Diagnosing a drop-off spike at a specific question
  • Managing quota fill rates and the composition drift they cause
  • Setting a defensible speeder threshold
  • Detecting straight-lining, duplicate identity and machine-generated open ends
  • Deciding whether to pause, amend or continue fieldwork
  • Monitoring qualitative and AI-moderated fieldwork for depth and abandonment

Typical inputs

What you give it.

Live field data with per-respondent status, timestamps and device metadata, The instrument with routing and expected paths, Quota structure and targets, Design assumptions for length, incidence, completion rate and source, A named decision owner who can authorise a pause or amendment, Previous wave or comparable field metrics (optional), Pilot or cognitive test findings (optional), Source variable on every respondent (optional)

Typical outputs

What you get back.

Pre-launch monitoring plan with timestamped thresholds, Soft launch report ending in release, release with fix, or hold, Field status dashboard with projected fill per quota cell, Drop-off diagnosis by question with the mechanism named, Quality flag summary with definitions, threshold dates and concentration by source and device, Composition monitor showing achieved composition by collection date, Intervention log with affected respondent ID ranges, Handover pack of flag variables and definitions for data preparation

Method coverage

What the skill works through.

  1. Why fieldwork is the last cheap moment to fix a study
  2. Writing the monitoring plan before launch, thresholds included
  3. Reading a soft launch: the seven checks, in order
  4. Diagnosing drop-off: what each pattern actually means
  5. Length in practice versus length in design
  6. Quota fill rates and the time confounding nobody notices
  7. Setting a speeder threshold you can defend
  8. Straight-lining, pattern responding and why one signal is not enough
  9. Open-end quality: duplicates, non-answers and machine-generated text
  10. Within-respondent contradiction checks
  11. Device, source and channel effects, and why to check them first
  12. Fraud and duplicate identity signals
  13. Pause, amend or continue: the decision framework and the log

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

What should I check in a soft launch?

Seven things, in this order because the earlier ones invalidate the later ones: did routing work; is incidence as expected; is median length as expected; where do people leave; do the open ends make sense (read every one at this stage); do the key measures have usable distributions; and do individual records cohere when read end to end. The output is a verdict of release, release with a named fix, or hold. A soft launch that produces no observations has not been read.

How do I set a defensible speeder threshold?

Three approaches are in common use and all are conventions rather than validated rules, so state which you used and why. An absolute floor derived from a reading-speed calculation over the instrument's actual word count is the most defensible because it is derived rather than borrowed. A relative threshold (a third or a half of the median) adapts to the instrument but becomes circular if the sample is already full of speeders. Section-level timing, catching people who are fast on the parts that require reading, is the most precise and needs question-level timing. Whichever you use, set it before seeing results and prefer flagging to deleting.

What does a drop-off spike at one question mean?

That question. Look for a required answer respondents cannot give, an option list missing their case, an intrusive question, a grid too wide for a phone, or a technical failure. Distinguish it from a spike immediately after a question, which usually means the question was answerable but unwelcome; a steady gradient, which is length and fatigue; a step at a section boundary, which is a transition problem; and early loss before the first substantive question, which points at the invitation or consent text.

Is straight-lining enough to exclude a respondent?

On its own, no. Some respondents legitimately hold the same view of every item, especially on short batteries with similar items. Corroborate it: use reverse-coded items, check whether the respondent straight-lines every battery or just one, and combine with timing. A grid completed faster than it could be read is a different finding from one completed slowly.

How can I tell if an open-ended answer was machine-generated?

Suspicious text tends to be fluent, well structured, longer than typical respondent writing, generic about the specific product or experience, and to answer the question more completely than a person would while containing no concrete detail. Two cautions matter. Fluency is not proof, and excluding articulate respondents biases the sample. And no detection method is reliable enough to justify silent removal, so flag for human review rather than deleting by rule. Duplication, non-answers and text copied from the question stem can be handled by rule.

Why do quotas filling at different speeds cause a problem?

Because it confounds subgroup with time. If younger respondents fill in two days and older respondents take three weeks, the two groups were collected in different periods, and anything that happened during fieldwork affects them unequally. The under-used fix is to throttle the fast cells so every cell runs across the whole window, rather than only boosting the slow ones.

How do I tell a bad respondent from a bad question?

Cross the signal by question, by source and by device. A signal that concentrates at a specific question, across all sources and devices, is an instrument problem. One that concentrates in one source, device or entry point, spread evenly across the instrument, is a sampling or mode problem. Only a signal that concentrates in individuals, on multiple independent indicators, across the whole instrument, is a respondent problem. Researchers reach for the third explanation first and it is the least often correct.

What is the strongest signal of fraud or misqualification?

A qualification rate far above the expected incidence. People do not become more common because you started asking about them. A screener that suddenly passes 40% of a population known to be 8% is not finding more of them. Other signals include implausible bursts of completes, identical open-end text across respondents, and screener answers conflicting with later answers. Report the signals and escalate the determination to a person, since rejecting a batch has contractual consequences.

When should I pause fieldwork?

When continuing would spend sample on data that will not be usable. Pausing costs field time and preserves sample, which is usually the better trade early in fieldwork and rarely worth it late. Before deciding, check how much sample is unspent: a defect found at 85% collected is a disclosure, not a fix.

Why does every mid-field change have to be logged?

Because an amendment partitions the sample into a pre-change group and a post-change group that are not exchangeable, and analysis needs to carry that partition. The log entry records the timestamp, what was observed, on what evidence, what was decided, by whom, what changed, and the respondent ID ranges before and after. An unlogged amendment is worse than the defect it corrected, because the defect at least remains visible in the data.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out