- Home
- Research skills
- Open-Ended Response Coding
Open-Ended Response Coding
Code hundreds or thousands of short survey verbatims into a defined frame, and report prevalence against the right base, with the non-answers counted.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
Open-ended survey questions produce the most honest material in a questionnaire and the most abused. A column of a few thousand fragments usually gets summarised into impressions with no counts behind it, machine-clustered into groups nobody can define, or coded to a frame that was never written down, so no two waves agree.
This skill supplies the discipline. Read a sample before building anything. Write every code with an inclusion rule, an exclusion rule and a boundary example, so a second coder would reach your answer. Decide multi-coding explicitly and accept that percentages will sum above 100. Compute nets on unique respondents rather than by adding up their member codes. Treat a large "other" bucket as proof the frame is wrong, not as a category. Classify the blanks, the refusals and the "nothing to add" responses instead of deleting them, because each is a different finding.
It is strict about two things that break open-end reporting: the denominator, where answered and asked can differ by a factor of two, and the meaning of the number, because an unprompted mention rate measures what was top of mind, not what people believe. Absence is not disagreement.
You get a documented code frame, a frequency table on both bases, a non-answer report, frame diagnostics, and a plain statement of what the question cannot tell you.
Best used for
- Reason-for-score follow-ups behind a rating or recommendation question
- Tracker open ends where wave-on-wave comparability must be preserved
- Unaided mention lists and unprompted awareness questions
- Exit, cancellation and complaint free text at volume
- Public consultation returns where every submission must be accounted for
- Explaining a quantitative result that the closed questions cannot explain
Typical inputs
What you give it.
Unfiltered open-ended response column with respondent identifiers, Exact question wording and the question that preceded it, Routing rules showing who was asked, Prior wave code frame with definitions (optional), The closed question the open end explains (optional), Respondent characteristics and weighting variables (optional), Client reporting taxonomy or existing categories (optional)
Typical outputs
What you get back.
Documented code frame with inclusion rules, exclusion rules and boundary examples, Net and super-code structure with unique-respondent computation rules, Frequency table on both answered and asked bases, Non-answer report separating blanks, refusals and explicit nothings, Frame diagnostics including other rate and per-code audit disagreement, Wave-on-wave comparison with frame changes logged, Salience statement explaining what the prevalence figures can and cannot mean, Illustrative verbatims verified against source with respondent IDs
Method coverage
What the skill works through.
- What open-ended coding is, and where it differs from thematic analysis
- Reading a sample before you build a frame
- Building a code frame bottom up
- Applying an existing frame, and why tracker frames should not be improved
- Writing code definitions: inclusion, exclusion, boundary example
- Multi-coding, and what it does to your percentages
- Nets and super-codes without double counting
- The "other" code as a diagnostic on the frame
- Coding the non-answer: blanks, refusals and "nothing"
- Choosing the base: answered, asked, or total sample
- What an unprompted mention rate actually measures
- Reviewing the frame before mass application, and auditing after
- When to use a different skill instead
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
How do I build a code frame for open-ended survey responses?
Read a random or stratified sample of 100 to 200 responses before writing a single code, so the frame is not shaped by whatever appeared first in the file. Then write codes at the smallest distinct level people are actually talking about, give each one a definition, an inclusion rule, an exclusion rule and a boundary example, and net upward for reporting. Pilot the frame on fresh responses before applying it to everything.
Should I report open ends as a percentage of everyone or a percentage of those who answered?
Show both, and state which one the commentary uses. On a question with a 45 percent answer rate the two figures differ by more than a factor of two, and a reader who assumes the wrong one will misread every finding on the page. This is the commonest reporting error in open-end work, and it happens in both directions.
Can one response have more than one code?
Usually yes, and it should. Most people say more than one thing. Multi-coding means the column will sum above 100 percent, which is fine as long as you state the convention and never convert it into a "share of mentions" table, because that quietly changes the denominator from people to codes.
What do I do with blank and "n/a" responses?
Classify them rather than deleting them. A blank, an explicit refusal, "nothing to add" and a junk string are four different findings. Deleting them silently changes the denominator without a log. "Nothing to add" after a satisfaction rating is often a substantive statement about the absence of problems and belongs in the coded data.
How big can the "other" code be before there is a problem?
Under 10 percent is normal residue. Between 10 and 15 percent the frame is missing something specific and you should find it. Above 15 percent the frame is wrong and should be rebuilt from a fresh sample. A large other bucket is not a category, it is the frame telling you it was built on the wrong material or at the wrong level of granularity.
How do I keep a code frame consistent across tracker waves?
Freeze it. Add new codes beneath the existing structure and log every addition, but do not restructure because you would have built it differently. A restructured frame converts frame change into apparent market change, and the trend is usually the client's entire reason for asking. If a rebuild is genuinely necessary, declare it as a break and run both frames in parallel for at least one wave so the discontinuity is measured.
Does a low mention rate mean people are not concerned about something?
No. An unprompted open end measures what was top of mind under that specific prompt. Nine percent mentioning waiting times does not mean 91 percent are content with waiting times: it means 91 percent raised something else. Absence is not disagreement. If you need the incidence of a belief, that requires a prompted closed question.
Is coding open ends the same as thematic analysis?
No, and the difference is the material. Thematic analysis works on rich, discursive text where meaning sits in how people explain themselves, and it builds themes with organising concepts, tested across the dataset. Open-end coding works on short text where the deliverable is a distribution: what was said, by how many, out of how many. Running thematic analysis on one-line answers produces a topic list dressed as an analysis.
Can AI code open-ended responses accurately?
It codes consistently, which is not the same as correctly. Two AI passes over the same data measure stability of application, not validity, and reporting an agreement coefficient over them dresses one up as the other. The controls that matter are sequencing and audit: a human approves the frame before it is applied at scale, and a sample of the coded output is re-coded blind afterwards, with disagreement reported per code rather than overall.
How do I compute net codes correctly?
On unique respondents. A respondent who gave two price-related codes counts once in the price net. Building a net by summing its member codes overstates it, sometimes badly, and the error grows with how freely you multi-coded. Nets are reporting objects, never coding objects, and nobody should ever code directly to one.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
