- Home
- Research skills
- Question Bias Detection
Question Bias Detection
Audit a questionnaire or guide as a hostile reviewer: a severity-rated defect list, each with mechanism, direction of effect and a full rewrite.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
Authors cannot audit their own instruments, and not because they lack skill. An author reads what they meant. They know which word was chosen carefully and which list was researched, and that knowledge fills in every gap a respondent will meet cold. Bias is also asymmetric in visibility: a leading adjective is caught by everyone, while a missing "not applicable", an unbalanced option list and a priming introduction are caught by almost nobody and do more damage.
This skill runs a systematic detection pass over an existing instrument, from the position of a reviewer rather than an author. It covers leading and loaded wording, double-barrelled questions, assumed premises, unbalanced scales, missing and unbalanced answer options, order and context effects, priming from earlier questions and from the introduction, acquiescence, social desirability, framing, absent "don't know" and "not applicable", and culturally or linguistically loaded terms. It also audits the bias no question-level review can find, the kind that sits in the objectives.
The output is a severity-rated defect register in which every entry names its mechanism, its likely direction of effect, whether it is recoverable at analysis, a full rewrite, and what the rewrite does not fix.
Best used for
- Reviewing a questionnaire before programming or translation
- Auditing a client-supplied or stakeholder-supplied instrument
- Checking an AI-generated draft before it is fielded
- Finding bias in option lists, missing answer options and scale balance
- Detecting order, priming and carryover effects between questions
- Finding the bias that sits in the objectives rather than the wording
- Diagnosing a fielded study that produced an implausible result
Typical inputs
What you give it.
The instrument in the form respondents will meet it, including options and order, Research objectives, and the survey introduction and invitation text, Target population and mode, Analysis plan, to judge recoverability (optional), Previous wave wording, where a trend is at stake (optional), Source of each substantive option list (optional), Language and translation status (optional)
Typical outputs
What you get back.
Audit scope statement naming which passes were and were not run, Field or do-not-field verdict with defect counts by severity, Objective and framing findings, held separately from question findings, Defect register with quoted text, mechanism, direction, severity, recoverability, rewrite and residual, Whole-instrument findings on order, priming, carryover and anchoring, Response-style findings on acquiescence, social desirability and translation equivalence, Drafted limitations text for defects that will not be fixed, Consolidated review points
Method coverage
What the skill works through.
- Why authors cannot audit their own instruments
- The bias in the objectives, which no question review will find
- Reading the instrument once, in order, as a respondent
- Stem defects: leading, loaded, double-barrelled, assumed premise
- Response frames: balance, exhaustiveness, and the options that are missing
- Order, priming, anchoring and carryover between questions
- Acquiescence, social desirability and response style
- Culturally and linguistically loaded terms
- Rating severity by consequence rather than conspicuousness
- Recoverable at analysis, or fatal
- Writing the rewrite, and naming what it does not fix
- Auditing your own rewrites
- Auditing an AI-generated instrument
- Auditing after fielding, when the data is the evidence
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
How do I know if a survey question is biased?
Work through four passes rather than reading for a feeling. Check the stem for leading and loaded wording, two ideas in one question, an assumed premise about what the respondent did or noticed, undefined quantifiers and jargon. Check the response frame for balance in both point count and label intensity, for options a respondent could legitimately double-select, for overlapping numeric ranges, and for the options that are missing. Check the whole instrument in order for priming, anchoring and contamination between questions. Then check for response-style effects such as acquiescence and social desirability. Most reviews only do the first pass, which is why most instruments field with their largest defects intact.
What is a double-barrelled question?
One that asks about two things and permits one answer, usually joined by "and" or "or": "how satisfied are you with the speed and accuracy of the service?". A respondent who is happy with one and not the other cannot answer it honestly, and the analyst cannot tell which half produced the score. It is generally unrecoverable at analysis, because no technique separates the two components after the fact. The fix is to split it, or, where length forbids, to keep the half that serves the decision and drop the other.
Why does a missing "don't know" option matter?
Because the respondents who did not know are now inside the substantive distribution and cannot be identified afterwards. Every missing non-substantive option has a direction: a missing "don't know" inflates whichever substantive answer is most available; a missing "not applicable" inflates the neutral or midpoint; a missing "none of these" on a multi-select inflates the nearest listed option; a missing "prefer not to say" on a sensitive question produces break-off or a false answer. The counterpart discipline matters too: adding "don't know" to an attitude question where everyone has a view manufactures non-response and rewards satisficing.
What is an unbalanced scale?
One where the positive and negative sides are not symmetric, in point count or in label intensity. "Excellent, very good, good, fair, poor" is the classic case: three positive points against one negative, with a mean pushed upward before anyone answers. It is partially recoverable, in that the bias applies to everyone so subgroup comparison and wave-on-wave movement can still be read, but the absolute level is uninterpretable and cannot be compared to any externally scaled measure.
How do earlier questions bias later ones?
Four mechanisms. Order contamination, where an aided list shown early destroys the ability to measure spontaneous awareness later, which is irreversible. Priming, where a block of questions establishes a frame the key question then sits inside. Carryover, where a respondent who has just committed to a position answers consistently with it whether or not that is what they think. And anchoring, where a number, a price or a range stated in one question sets the scale for the answers to the next. None of these is visible when reading question by question, which is why the instrument has to be read once end to end, in order, before anything is annotated.
Can bias be in the objectives rather than the questions?
Yes, and it is the most consequential kind, because every question can be faultless and the study still cannot return the finding that matters. "Understand why customers value the new service" cannot produce the finding that they do not. "Identify the barriers to adoption" presupposes that adoption is desirable and that the obstacle is on the customer's side. The test is to rewrite each objective symmetrically and see whether the instrument still makes sense; frequently half of it does not, and the gap is the missing finding. This cannot be fixed by rewriting a question and has to go back to whoever agreed the brief.
How should I rate the severity of the defects I find?
By consequence, not by how obviously wrong the question looks. A badly worded question nobody will report is minor. A subtly unbalanced scale on the primary measure is critical. Two things push severity up: whether the question feeds something that will be reported, and whether the defect is recoverable at analysis. A reviewer who marks everything as important gets everything ignored, so a stated rubric with a small number of critical findings is far more likely to result in changes than a long undifferentiated list.
Should a bias audit include rewrites?
Always, in full text. A flag is a complaint; a rewrite is a deliverable. Reviewers who supply replacement wording get their changes made, and reviewers who supply adjectives get a defensive meeting. Two disciplines go with it: state in one line what the rewrite does not fix, because a neutrally reworded question still sits after the block that primed it, and then audit the rewrites themselves, because corrections introduce new defects at a high enough rate that skipping that step is a known way to make an instrument worse.
How do I review an AI-generated questionnaire?
Target the places generation is weakest rather than reading it front to back for tone. Generated drafts are fluent, so stem-level defects are rarer than in human drafts. They fail most consistently on option lists (frequently invented, frequently containing every positive attribute and one token negative), on scale label symmetry, on missing non-substantive options, on acquiescence load because agree/disagree batteries are easy to generate symmetrically, and on length, because generated instruments have nothing removed. Ask for the provenance of every substantive list: an unsourced brand, feature or competitor list is a defect of unknown direction, not a neutral default.
What if the biased question is part of a tracker and cannot be changed?
Then the right output is a disclosure, not a change request. Log the defect, state its direction, and note what remains readable: an unbalanced scale that has run for six waves still supports wave-on-wave comparison, because the bias applies equally every time, even though the absolute level cannot be compared to anything external. Present the options and their costs (change and re-baseline with a parallel run, or retain and disclose) and leave the decision with the researcher, because trading measurement quality against a published commitment is a business judgement, not an audit finding.
Can an instrument ever be certified unbiased?
No, and an audit that says so has overstated its product. Every instrument makes choices that push answers in some direction, and some of those choices cannot be avoided: the sponsor is inferable, the mode has effects, the topic has a socially preferred answer. The honest output is what was checked, what was found, what was fixed and what remains, with the direction of each residual stated, so that the eventual report's limitations section can be written from the audit rather than invented at the end.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
