07.01Qualitative AnalysisAvailable

Thematic Analysis

Turn interview and focus group transcripts into tested, evidenced themes, with honest prevalence, preserved contradictions and every quote traceable to a participant.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Thematic analysis is the workhorse of qualitative research and the easiest thing to do badly. Read the transcripts, notice what recurs, write it up, attach three quotes: the result reads like analysis and is closer to impression. It cannot say how many participants held a view, whether anyone held the opposite, or whether the theme survives transcripts it was not built on.

This skill encodes the professional method. It defaults to reflexive thematic analysis with a codebook discipline borrowed from framework analysis, and says where grounded theory or framework analysis fits better. It runs familiarisation, initial coding, code frame construction with inclusion and exclusion rules, theme testing across the whole dataset, and deliberate sweeps for contradiction and divergence.

Two things it refuses to do: reduce qualitative material to word frequency, and resolve ambivalence into a cleaner statement than the data supports. It separates how often a theme appeared from how much it matters, so a view held by five people can be the study's most important finding, with the reasoning shown.

You get a documented code frame, theme records with definitions and boundaries, prevalence counted in participants, a divergence register, an honest saturation statement, and quotes that keep their participant IDs.

Best used for

  • Analysing 20 to 40 depth interviews as a set
  • Cross-group analysis of focus groups
  • Qualitative phases of mixed-method studies
  • Work where contradiction and ambivalence are part of the finding
  • Analyses that must withstand a research director or examiner asking "who said it"
  • Building a reusable code frame for a research repository

Typical inputs

What you give it.

Depth interview transcripts with participant identifiers, Focus group transcripts with speaker labels, Long-form open-ended responses, Diary, community or longitudinal qualitative entries, Ethnographic or accompanied-shop field notes, Discussion guide or question set used in fieldwork, Research objectives, Participant characteristics file (optional), Prior code frame from an earlier wave (optional)

Typical outputs

What you get back.

Documented code frame with definitions, inclusion and exclusion rules and boundary examples, Fully coded dataset with participant IDs and verbatim spans retained, Coding stability check with per-code disagreement rates, Theme records with definition, organising concept, boundary and evidence, Prevalence table in participants with stated denominators, Divergence and contradiction register, Saturation statement, or an explicit declination to claim saturation, Importance assessment reported separately from prevalence, A section stating what could not be established

Method coverage

What the skill works through.

  1. What thematic analysis is, and what it is not
  2. Choosing your tradition: reflexive TA, framework analysis, grounded theory
  3. Inductive, deductive and hybrid coding
  4. Building a code frame that holds: definitions, inclusion and exclusion rules, boundary examples
  5. Coding the full dataset, and checking consistency when AI does the coding
  6. Developing candidate themes and testing them against the whole corpus
  7. Contradiction, ambivalence and minority perspectives
  8. Frequent versus important: why a theme raised by five can outrank one raised by fifty
  9. Reporting prevalence honestly
  10. Saturation, and when you cannot claim it
  11. Naming themes: a label, not a finding
  12. Quote integrity and participant traceability
  13. Where a human researcher must review
  14. When to use a different skill instead

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How do I use AI for thematic analysis?

Use it for the parts that reward consistency and scale: initial coding across a whole corpus, applying an approved code frame to every transcript identically, tracking prevalence by participant, and finding passages that contradict a theme. Do not use it to decide what a theme means, to resolve ambiguity, or to interpret cultural context. The critical control is sequencing: a human reviews the code frame before it is applied at scale, because a definitional error applied to thirty transcripts is thirty errors.

Can AI code qualitative data reliably?

It can code consistently, which is not the same thing. Two AI passes over the same material measure stability, not validity: they show the frame was applied the same way twice, not that it was applied correctly. Report it as stability, adjudicate the disagreements with a human, and fix the code definitions rather than the individual instances.

What is the difference between reflexive thematic analysis, framework analysis and grounded theory?

Reflexive thematic analysis treats themes as constructed by an analyst making visible decisions, and is the best default for most applied research. Framework analysis suits fixed objectives, cross-case comparison and clients who need a matrix. Grounded theory is for building theory and requires theoretical sampling and constant comparison, which most commercial timelines do not permit. Say which you used and why.

Should I code inductively or deductively?

Hybrid, usually. A deductive spine drawn from the research objectives ensures the analysis answers the questions the study was commissioned to answer. Inductive freedom below it lets the material say things nobody anticipated. Mark which codes came from where, so a reader can see how much of the frame was imposed.

How many participants make a theme?

There is no threshold, and any number offered as one is arbitrary. Report the count and the denominator and let the reader judge. Prevalence tells you how widely a view is held. It does not tell you how much it matters, and those two judgements belong in separate columns.

How do I know if a theme is important rather than just frequent?

Assess it against the decision the research informs, the consequence if it is true, whether it explains other themes, whether the people who hold it occupy a structurally significant position, whether anyone could act on it, and whether it is already known. A theme raised by five recent defectors often outranks one raised by everyone, because everyone is describing the category and the five are describing the failure.

What is saturation, and can I claim it?

Code saturation means no new codes are appearing. Meaning saturation, which matters more, means no new nuance is appearing within existing codes. You can only assess either after the fact, by tracking when new codes appeared. If new codes were still arriving in your final transcripts, you did not reach saturation, and a sample size on its own never demonstrates one.

How do I stop AI inventing or smoothing quotes?

Keep the participant ID attached to every extract from transcript to report, verify every client-facing quote word for word against source, and allow only the standard edits: filler removal, marked elision, bracketed clarification, flagged transcription corrections. Never merge two speakers, never present a paraphrase as verbatim, and check that a quote attributed to a segment came from someone in it. A quote that arrives without an ID cannot be checked, and an uncheckable quote is functionally a fabricated one.

Why does AI keep flattening contradictions in my qualitative data?

Because language models optimise for coherence, and ambivalence reads as noise to be cleaned up. This is the single most damaging failure in AI-assisted qualitative work, because the ambivalence is frequently the finding. Run contradiction as a separate deliberate pass, and write the tension into the theme definition rather than choosing a side.

When should I not use thematic analysis?

When the material is too thin to carry meaning (use open-ended response coding), when there are more than roughly 5,000 items and genuine familiarisation is impossible (use large-scale text analytics), when the unit of interest is one person (use transcript analysis), or when the real question is how many, which a qualitative sample cannot answer.

Do I still need a researcher if AI does the analysis?

Yes, at three specific points: reviewing the code frame before it is applied at scale, adjudicating ambiguous passages where two readings are both defensible, and interpreting cultural or linguistic context. These are positional limits rather than capability ones, and the output marks them explicitly rather than leaving them for the reader to find.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out