12.06Report Design and CompilationAvailable

Research Report QA

The last gate before a report leaves: seven structured passes over numbers, evidence, language, consistency and completeness, ending in a defect log and a judgement.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Most research reports are checked in the way that catches the fewest defects: one person reads the document once, for sense and typos, on the last afternoon, having already read it three times.

That process misses the failures that matter, and it misses them systematically. A number gets verified against the previous draft rather than against source, so the error is confirmed rather than found, and by version four it has been checked three times. A quote reads well because it was tidied during a rewrite, and nobody opens the transcript. A difference is called "significantly higher" with no test behind it. A finding that says "appears to be associated with" in the analysis says "drives" in the executive summary, having gained certainty at every restatement. And the limitation that qualifies the headline sits on page 60.

This skill replaces the read-through with seven defined passes, each looking for one class of defect across the whole document including the summary, chart titles and appendix: numbers, evidence, language, consistency, completeness, presentation, and judgement. It runs the significance check, the causal-language check and the confidence-calibration check as explicit separate passes, because folding them into a general read is how they get missed.

The output is not comments. It is a defect log with severity, location and required fix, and a pass or fail decision that holds on the day it is inconvenient.

Best used for

  • Reports about to go to a client, regulator or public audience
  • Decks built by several people and never checked as one document
  • Deliverables whose claims will be interrogated line by line
  • Tracking waves publishing trend or comparability claims
  • Reports inherited late where defensibility must be established
  • Any deliverable carrying a named researcher's sign-off

Typical inputs

What you give it.

Frozen final draft, including summary, appendix, charts and executive artefact, Source material behind every claim, including analysis tables and transcripts, Evidence map linking claims to sources, Agreed objectives, Questionnaire or discussion guide as fielded, with routing, Conventions register from compilation, Raw dataset where available, Statistical test outputs, Previous wave report and questionnaire, Accessibility specification where a standard is claimed

Typical outputs

What you get back.

Defect log with location, severity, required fix and source checked against, Verification record by item type, including items that could not be verified, Significance, causal, confidence-calibration and overgeneralisation pass results, Worked disclosure checklist against K4 requirements, Pass, pass with conditions, or fail judgement with the reason, Named holder of the release decision, Escalations to methodological review where a defect is not documentary

Method coverage

What the skill works through.

  1. Why a read-through misses the defects that matter
  2. Freezing the draft and building the item inventory
  3. The numbers pass: checking against source, not against the previous draft
  4. Checking every base and denominator
  5. The evidence pass: quotes, citations and recommendation traceability
  6. The significance-language check
  7. The causal-language check
  8. The confidence-calibration check
  9. Internal consistency and the same figure in three places
  10. Summary versus body: how claims strengthen in transit
  11. Caveat travel: limitations that stayed in the appendix
  12. Completeness: objectives, contradicting evidence, required disclosures
  13. Presentation and accessibility
  14. Severity, the defect log, and the pass or fail judgement
  15. What this checks, and what a research quality review checks instead

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How do I quality check a research report before it goes out?

Not by reading it once for everything. Run separate passes, each looking for one class of defect across the whole document: numbers, evidence, language, consistency, completeness, presentation. Then classify what you find by severity and issue a judgement. Attention cannot hold seven checking frames at once, which is why a single read reliably finds fluency problems and misses evidence problems. Every pass covers the executive summary, chart titles, headings, callouts and appendix, because defects concentrate in the material written last and read least.

Why should I check numbers against source rather than against the previous draft?

Because checking against the previous draft is how an error becomes established. Each version confirms the one before it, and by the fourth draft a wrong figure has been checked three times and everyone is confident in it. Open the analysis file, find the figure, and confirm the value, the base description, the base size, the filter, the weighting state and the question reference. Recompute anything that was typed by hand, which in practice means every number in a chart title, a callout, a headline or the summary.

How do I check for significance claims that were never tested?

Run it as an explicit pass. Search the document for "significant", "significantly", "notably", "markedly", "clearly", "substantially higher" and equivalents, including in headings and chart titles. For each, confirm a test was run and that the test, result and threshold are reported. Where no test was run, the wording changes to an observed comparison with both bases shown, or the claim goes. Check the reverse defect too: a tested and significant difference reported without its test understates what the study actually established.

How do I check for causal language a study cannot support?

Search for "drives", "leads to", "causes", "results in", "the impact of", "improving X will", "because" and "due to", and look hardest in chart titles and headings, where causal readings survive longest. For each, ask whether the design licenses a causal claim: a randomised experiment, a valid quasi-experimental design with its assumptions stated, or a longitudinal design with temporal order established. Where it does not, rewrite to association. Where a modelling output is conventionally called a "driver", the document should state once that these are statistical associations whose causal direction was not established.

Why does the executive summary so often overstate the body?

Because a claim loses a qualifier at each restatement, and each restatement is checked against the one before it rather than against the analysis. The fix is a specific check: compare every summary claim to the body claim it summarises, and compare both to the analysis output directly. Look for a hedge removed, a base dropped, a segment claim generalised to the whole sample, an association become a cause, and a moderate reading become a flat statement. This is the single most productive check in the whole pass.

Do I need to check every quote?

Yes, every quote, every time, word for word against transcript, with the participant identifier confirmed and the participant confirmed as belonging to the segment the quote is attributed to. Spot-checking selects against the quotes that need checking, because a tidied quote reads better than an untidied one and so is less likely to be chosen. A quote that cannot be located in a transcript is a critical defect, not a query.

What counts as a critical defect in a research report?

Anything where the document asserts something its evidence does not support, or omits something a reader needs in order to act correctly: a wrong number, an unverifiable quote, a recommendation with no traceable finding, a significance or causal claim the design does not license, a missing base where the base changes the reading, a summary claim stronger than the analysis supports, a material limitation absent from the document, or a required disclosure omitted. Severity is judged by consequence to the reader, not by how hard the fix is: a wrong number corrected in ten seconds is critical, and a terminology inconsistency across forty pages is minor.

Should QA produce comments or a decision?

A decision. A log of comments without a judgement leaves the release call to whoever is most tired at the end of the project. Fail while any critical defect is open, regardless of the deadline: delivering a document you know asserts something unsupported is a worse position than delivering late. Fail on a pattern of major defects even where each is individually tolerable, because fifteen inconsistencies indicate a systemic problem that spot fixes will not address.

Is report QA the same as a research quality review?

No, and the distinction matters. Report QA checks the document against its own evidence: does it say what its sources support, consistently, completely and defensibly. A research quality review checks the study: whether the design was appropriate, the sample adequate, the instrument valid, and the analysis the right one. A report can pass one and fail the other in both directions. Where QA reveals that a defect is methodological rather than documentary, it escalates rather than logging it as a document defect, and the report does not leave in the meantime.

What if I have no time for a full QA pass?

Do not sample randomly. Run the highest-yield subset in order: every number in the executive summary and every chart title, every recommendation traced backwards to findings, every quote, and the significance and causal language passes. That takes a fraction of the time and catches most critical defects, because critical defects concentrate in exactly those places. Then state in the log which passes were not run, so nobody reads a partial check as a full one.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out