05.02Quantitative AnalysisAvailable

Statistical Testing

Run the right significance test, report effect size and confidence intervals with the p-value, and tell a real difference from a merely significant one.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Two opposite failures dominate quantitative research. The first is claiming a difference that is not there: 44% against 39% on bases of 120 and 140, written up as fact. The second is treating significance as though it meant importance: on a sample of 8,000, a two-point gap tests significant and is reported with the weight of a twenty-point one. Underneath both sits cross-tab work, where a banner generates thousands of tests and the analyst reports the ones that came up.

This skill prevents all three. It selects the test from the design and the data type, covering proportions, means, ordinal and paired data and their non-parametric alternatives, plus what to do when an assumption fails. It requires an effect size and confidence interval alongside every p-value, separating "could this be noise?" from "is this big enough to matter?". It sets out the false-positive arithmetic behind multiple comparisons, the correction methods, and the pre-specification discipline that beats them. It teaches how to read a null result using power, and names the cases where no test should be run at all: census data, overlapping subgroups, small bases.

It produces a significance test log and a reporting standard for each claimed difference.

Best used for

  • Deciding whether a subgroup difference in a cross-tab is reportable
  • Wave-on-wave testing in a tracker
  • Comparing a test cell against a control
  • Interpreting a non-significant result correctly
  • Setting the testing rules for a tabulation job before it runs
  • Auditing a draft report for unsupported claims of difference
  • Controlling false positives across a large banner

Typical inputs

What you give it.

The specific comparison, stated as a question with a base for each group, Data type and design (proportion, mean, ordinal, count; paired or independent), Unweighted and weighted base sizes for every group, Sampling approach (probability, quota, panel, convenience, unknown), Analysis plan with pre-specified comparisons (optional), Agreed materiality threshold (optional), Respondent-level data rather than a tabulation (optional), Design effect, effective base or replicate weights (optional)

Typical outputs

What you get back.

Significance test log covering every test run, including null results, Reporting line for each claimed difference, with bases, test, threshold, effect size and interval, Effect sizes (percentage-point difference, Cohen's d and h, eta-squared, Cramér's V, odds ratio), Confidence intervals on differences, Multiplicity statement with the test count and correction approach, Minimum detectable effect for every null result, Testing note for the deliverable, Consolidated review points on materiality judgements

Method coverage

What the skill works through.

  1. The two opposite failures: false differences and trivial ones
  2. Before you test: does the data license inference at all?
  3. Choosing the test from the design and the data type
  4. Proportions, means, ordinal data and their non-parametric alternatives
  5. Paired versus independent samples, and why this is the commonest error
  6. Assumptions, and what to do when they fail
  7. Statistically significant is not the same as practically meaningful
  8. Effect size: Cohen's d and h, eta-squared, Cramér's V, odds ratios
  9. Why confidence intervals usually tell you more than p-values
  10. The multiple comparisons problem in cross-tab work
  11. Bonferroni, Holm, false discovery rate, and when correction is appropriate
  12. Pre-specified comparisons: the better answer
  13. Power, and how to read a null result
  14. When not to test at all
  15. What must accompany every claimed difference

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How do I know whether a difference between two survey subgroups is real?

Run the test that matches the design, then look at three things together rather than one. The p-value tells you whether the gap is larger than sampling noise. The confidence interval on the difference tells you how large the gap plausibly is. The effect size tells you how large it is in units a reader can judge. A difference reported without all three, and without both base sizes, is not reportable.

What is the difference between statistical significance and practical significance?

They answer different questions. Statistical significance asks whether a difference could be noise. Practical significance asks whether it is big enough to matter. A large sample makes trivial differences significant and a small sample hides real ones, so significance on its own tells you as much about your sample size as about the world. Both must be reported, and the practical judgement needs a materiality threshold agreed before the analysis, not argued afterwards.

Which statistical test should I use?

It follows from two things: what the measure is, and whether the same respondents appear in both groups. Two independent groups on a proportion take a two-proportion z-test; on a mean, Welch's t-test; on ordinal or heavily skewed data, Mann-Whitney U. Three or more independent groups take a chi-square, an ANOVA or a Kruskal-Wallis, each followed by targeted comparisons. The same respondents measured twice take McNemar's test, a paired t-test or a Wilcoxon signed-rank test. Running a paired design through an independent test is the commonest selection error in tracker and pre/post work.

What is a good effect size?

There is no universal answer, and the conventional small, medium and large labels were offered as a last resort for fields with no established scale. A standardised difference of 0.2 in a mature category where nothing has moved for five years can be the largest effect anyone will find. Calibrate against the measure's own historical variation, the size of effects comparable interventions have produced, and the difference the business would actually act on.

Why do so many differences come up significant in a cross-tab?

Because a banner runs an enormous number of tests. Twelve columns produce sixty-six pairwise comparisons per row; across sixty rows that is nearly four thousand tests, of which roughly two hundred will flag at a 5% threshold even if nothing differs anywhere. Reading the flagged cells off the table and writing them up reports noise in a formal-looking way. Count the tests before running them, and either correct within a defined family or label the output as exploratory screening.

When should I apply a Bonferroni correction?

When you are screening a family of comparisons and any one of them might be reported as a finding. Bonferroni is simple and valid but conservative; Holm controls the same error rate and is uniformly more powerful, so there is little reason to prefer Bonferroni once Holm is available. Benjamini-Hochberg controls the false discovery rate and is usually the better choice for exploratory screening. None of these converts a fishing expedition into a confirmatory result. A small set of comparisons pre-specified before the data was seen is the real solution.

Does a non-significant result mean there is no difference?

No. It means the study could not distinguish the observed gap from zero, which is exactly what a real difference looks like when the sample is too small to detect it. Before writing anything about a null, work out the minimum detectable effect. For two proportions near 50%, roughly 20 percentage points at n=100 per group, 14 at n=200 and 9 at n=500. Then say what the study could and could not have seen. Establishing that two things are genuinely equivalent needs an equivalence margin set in advance, not a failed test.

Can you run significance tests on a non-probability sample?

The theoretical basis for the p-value assumes random selection from a population, and most commercial online samples are not obtained that way. Practice is genuinely split between not testing at all and testing as a rough filter for whether a gap exceeds ordinary random variation. The defensible position is to test if you choose, describe the result as a noise filter rather than inference to a population, and report no margin of error.

Why can I not test my weighted data the normal way?

Because a standard test treats weighted counts as though the weights were extra respondents, and that manufactures significance. Calculate the effective base first, then use it in every test. A design effect of 1.6 means a nominal base of 800 behaves like 500, and the difference is enough to flip results either way.

Can AI run significance testing reliably?

It can run the arithmetic consistently, which is a genuine advantage. The risks are specific: producing a plausible p-value with no calculation behind it, choosing a test from the shape of a table without knowing whether the respondents are the same people, reporting only the comparisons that came up positive, and attaching mechanical effect-size labels with no reference to the measure or the decision. Used properly, AI should be required to produce a full test log including the nulls, state the number of tests run, and refuse to test where the bases, the weighting or the sample design do not support it.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out