- Home
- Research skills
- Quote and Evidence Extraction
Quote and Evidence Extraction
Select verbatim evidence for representativeness rather than eloquence, verify every quote against source, and never manufacture one where none exists.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
This is where fabricated quotes enter research reports. Not through malice, but through a chain of small reasonable-looking steps: a theme needs evidence, a striking phrase surfaces from memory, a near-match is tidied to fit the slide, two participants' words are combined because neither said the whole thing, and a sentence nobody spoke arrives in an executive summary carrying more persuasive weight than any statistic in the deck.
This skill is the control. It starts from the claim, not the quote. It builds the candidate pool from coded material rather than recall, screens for typicality before quality, and applies a coverage budget so a twenty-participant study does not become four people quoted twelve times. Every quote is matched to the specific claim it sits under, because a genuine quote under an adjacent claim evidences nothing while reading as though it does. Every quote is verified word for word against source, which is required and not a quality-assurance nicety. Editing is limited to the permitted operations and the deliverable carries a stated convention. Attributions are screened for re-identification, because a job title plus a market is frequently one person.
Where a real theme has no quotable verbatim, this skill reports the theme with its prevalence and a labelled analyst summary. It does not write the missing quote.
Best used for
- Selecting supporting verbatims for each theme in a qualitative report
- Choosing the quote that opens a section or presentation
- Building an evidence appendix that withstands a client audit
- Anonymising quotes for small or specialist samples
- Auditing an existing deck's quotes back to source
- Representing a dissenting minority without inflating it
- External-facing, published or regulatory outputs where every word is scrutinised
Typical inputs
What you give it.
Full searchable source text (transcripts or coded open ends), The written claims the evidence must support, Coded or themed material with participant identifiers, Participant characteristics as they will be printed, Theme prevalence data (optional), Audio or timecodes for verification (optional), Consent wording covering verbatim reproduction (optional), Report or slide layout constraints (optional)
Typical outputs
What you get back.
Evidence register with claim, quote, source location, verification status and edits, Quote convention statement for the deliverable, Coverage summary showing distinct participants and per-participant concentration, Typicality ratings distinguishing modal from outlier expressions, Minority-view quotes with base attached in the same visual unit, Re-identification screen result on every printed attribution, List of claims with no adequate verbatim support and how they are reported instead, Rejected-quotes list with selection reasons
Method coverage
What the skill works through.
- Why quote selection is where research reports fail
- Starting from the claim, not the quote
- Building a candidate pool you did not remember
- Representativeness versus eloquence: the articulate-respondent bias
- Illustrative or merely memorable
- Matching the quote to the claim it evidences
- Coverage: how many distinct participants your report actually quotes
- Verifying every quote against source
- Permitted edits and the convention statement your report must carry
- Attribution, anonymisation and re-identification in small samples
- Quoting a minority view without giving it false weight
- What to do when no good quote exists
- When to use a different skill instead
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
How do I choose which quotes to use in a research report?
Start from the claim each quote has to evidence, written out as it will appear in the report. Pull every coded passage relevant to that claim, read them as a set to establish what the ordinary expression of the theme sounds like, then select for typicality before quality. Choosing a quote first and writing the claim around it inverts the logic of the report and turns a memorable phrase into a finding.
How do I stop AI fabricating or smoothing quotes?
Require that every quote is located in the source text and compared character by character, in the same session it is used. Consistency with the transcript is not verification against it. Allow only four edits: filler removal, marked elision, bracketed clarification, and flagged correction of an obvious transcription error. Never permit merging two speakers, and treat any quote that captures a theme unusually completely as a suspect rather than a prize.
What edits to a quote are acceptable?
Removing filler where meaning is unchanged, marked elision, square-bracketed clarification of a referent, and flagged correction of an obvious transcription error. Not grammar, not tense, not word order, and not the removal of a hedge that weakens the point. Whatever you do must be stated once, visibly, in the deliverable as a quote convention. If a quote needs more work than that to make the point, the quote does not make the point.
Why are my quotes always from the same few participants?
Because selection on how well something reads systematically favours the fluent, and fluency correlates with education, confidence and familiarity with being asked one's opinion, none of which the study is measuring. The correction is procedural: set a coverage budget before selecting, count distinct participants across the whole report rather than per theme, and deliberately choose less well expressed evidence from other people when the same voice keeps winning.
How do I anonymise quotes properly?
Anonymity is a property of the descriptor set, not of the identifier. A participant number protects nobody when the attribution reads "procurement director, mid-size manufacturer, northern region", which is often one person. Build the attribution string exactly as it will be printed and ask how many people in the source population match it. Then reduce descriptors, band characteristics, or move to a group-attributed paraphrase.
What if there is no good quote for a real theme?
Report the theme with its prevalence and describe it in the analyst's voice, clearly labelled as your summary rather than a participant's words. This is common and it is not a failure of the analysis: some themes are expressed in fragments and half-sentences by everybody. Composing a sentence that captures what participants meant, or merging two people's phrasing, is fabrication regardless of how faithfully it reflects the material.
How do I quote a minority view without exaggerating it?
Put the base in the same visual unit as the quote, not in a footnote and not on the previous slide. State the prevalence in the framing sentence before the quote is read. And do not select the most forceful expression of the minority view, because that compounds the weighting problem. A reader who saw only that slide should correctly estimate how many people held the view.
Does every quote really need verifying against the source?
Yes, including the ones that seem obviously fine. A quote is the most persuasive object in a research report and the least checkable once it is published, and the check will not happen downstream. Verify the quote you like best first: perfection is a symptom, because real speech is redundant, hedged and slightly off the point.
What is the difference between an illustrative quote and a memorable one?
An illustrative quote lets a reader recognise the theme in new material. A memorable quote is one they will repeat. These are different properties that sometimes coincide, and the memorable one travels further, which means it does more damage when it is atypical. Choosing for memorability is a legitimate communication decision if you make it explicitly and place a typical quote alongside.
How do I check the quotes in a report someone else wrote?
Audit rather than reselect: locate every quote in the source, check attribution against the participant's own record rather than the segment the claim is about, check for splicing across two locations, and run the re-identification screen on the printed attribution string. Report anything unverifiable as unverifiable rather than deleting it quietly, because its presence tells you something about the process that produced the document.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
