- Home
- Research skills
- AI Output Verification
AI Output Verification
Audit AI-produced research against the sources it was given, using checks built for how AI fails rather than how people do.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
AI research output fails differently from human research output, and a reviewer applying human-error instincts looks in the wrong places. A tired analyst produces work that is visibly rough. An AI system produces work that is uniformly fluent, internally consistent, correctly formatted and wrong in ways that leave no surface trace. The output looks most finished exactly where it is most invented, because a fabricated specific has no messy provenance to complicate it.
This skill audits against that failure profile directly: fabricated specifics, plausible citations, restatement presented as synthesis, silent gap-filling, ambiguity resolved toward the more coherent reading, contradictory material quietly dropped, and uniform confidence regardless of evidence strength.
The procedure establishes what the AI actually received before checking what it produced, recomputes numbers against source, verifies every quote and every citation without sampling, traces claims back through the evidence chain, and then does the step most verifications skip: audits for what is absent, working from the inputs rather than from the output's own structure.
It ends in a trust judgement of four forms, the last of which is "cannot be relied on", and a disclosure record of what a human actually checked.
Best used for
- Auditing AI-produced coding, classification or theme development
- Checking an AI-drafted analysis or report section before client delivery
- Verifying an AI synthesis of transcripts, open ends or documents
- Establishing whether AI-generated insights are evidenced or merely coherent
- Producing the human-verification record required for disclosure
- Auditing an inherited output of uncertain provenance
Typical inputs
What you give it.
The complete AI output in the form it will be used, The full input set the AI was given, The instruction or prompt and the processing sequence, Statement of which parts are AI-produced and which are human-written, Model or system version and settings, Intermediate outputs from each step in a chain, Human-coded comparison sample for classification work, The analysis plan or code frame the output was to apply
Typical outputs
What you get back.
Verification scope statement with the coverage rule applied, Trust judgement in one of four forms, Verification log by claim type with source checked and result, Absence audit listing unrepresented inputs, dropped contradictions and resolved ambiguities, Calibration assessment against K3 confidence levels, Corrections required with consequential edits, Disclosure record of what the AI did and what a human verified
Method coverage
What the skill works through.
- How AI research output fails, and why it looks different from human error
- Establishing the verification boundary
- Checking whether the AI actually received all its inputs
- Predicting where this particular output will fail
- Verifying numbers: what is exhaustive and what can be sampled
- Verifying every quote, without exception
- Verifying citations at total coverage
- Tracing claims back through the evidence chain
- Significance language, causal drift and confidence calibration
- The absence audit: finding silent gap-filling
- Dropped contradictions and over-resolved ambiguity
- Adversarial prompting, and how to read the response
- The trust judgement, and when an output cannot be repaired
- The disclosure record that makes AI use defensible
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
How do I know whether an AI-generated insight is real?
Trace it back to source. Anything specific in the output that cannot be located in the input set, or derived from it by a stated operation, is a candidate fabrication carrying the burden of proof. For an insight specifically, apply two tests: does it explain why rather than restate a finding at a higher altitude, and does a named piece of source material sit underneath it. Restatement presented as synthesis is one of the most common AI failures and it reads exactly like insight.
What should I check first in an AI-produced analysis?
Whether the AI actually received all its inputs. Truncation, failed file loads and context limits produce outputs that describe a subset with the confidence of a census, and this is invisible from the output because the output never mentions what it did not see. Check participant identifiers, document coverage and the arithmetic of prevalence counts against the number of items you supplied. This one check can invalidate an entire synthesis in five minutes.
Can I verify an AI output by sampling?
Partly. Some classes are never sampled: every quote, every citation, every base size, every prevalence count, every number in the executive summary, headlines, chart titles and recommendations. Everything else can be sampled, stratified across sections and claim types rather than taken in document order. And there is an escalation rule: any fabrication found in a sampled class converts that class to exhaustive coverage, because one invented specific is evidence about the process that produced all the others.
How do I check for things the AI left out?
Work from the inputs, never from the output's structure. Build an inventory of what went in (every question, participant, document, code) and check each against the output. Then run three searches: find the strongest evidence in the source that cuts against the main conclusion and see whether it appears; find passages that support two readings and see whether the ambiguity was acknowledged or silently resolved; and check whether the output says anywhere that something could not be established. An output with no gaps is describing a research project that does not exist.
Why does AI output sound equally confident about strong and weak findings?
Because confidence in generated text is a property of register rather than of evidence assessment. The practical consequence is that a reader cannot distinguish a claim resting on 1,200 responses from one resting on three interviews. Read the output for its confidence register alone, ignoring content, and ask whether confidence varies at all. Uniform confidence is itself a finding, and it means no calibration was performed, so correcting individual claims does not fix it.
Does challenging an AI's output tell you whether it is right?
It generates leads, not verdicts. Ask which participant said something and where in the transcript: a grounded claim answers with an identifier that checks out, an ungrounded one produces a new plausible identifier, which is a second fabrication and a strong signal. Ask for the evidence that cuts against the conclusion: genuine analysis produces specific contrary material, gap-filled analysis produces a generic caveat. Then verify every lead against source, because folding is not proof of error and defending is not proof of correctness.
When can an AI output not be repaired by correcting the errors?
When the verified parts cannot be separated from the unverified ones. A fabricated citation is repaired by deletion. A fabricated prevalence count is not, because it means the analysis was not performed on the data and every other count is now unknown. Similarly, if a central theme was built by omitting the participants who contradicted it, correcting the counts does not restore the theme. That is a case for redoing the work, not for editing it.
How do I verify AI coding or classification at scale?
With a measured agreement rate rather than item-level checking. A human independently codes a random sample, and agreement is computed per code rather than overall, because rare and boundary categories are where accuracy collapses while the aggregate stays high. Stratify the sample to over-represent rare codes. Report the per-code rates, and treat the worst rate among codes that carry a finding as the number a client should rely on.
What has to be disclosed when AI was used in research?
What the AI did, what a human verified, at what coverage, what was found, and who performed the verification. That record, dated and named, is what makes AI use defensible, because nobody downstream can assess output quality directly. Disclosure is increasingly a client and regulatory requirement, and the standing obligation is not to present AI-assisted work as though it were not.
Is verifying an AI output the same as reviewing the research?
No. Verification checks whether an output faithfully represents the inputs it was given. Quality review checks whether the underlying study could support the claims at all. An output can be perfectly faithful to a study that should never have been run that way. Run verification first, then bring its report into the quality review as an input.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
