06.04Specialist and Advanced AnalysisAvailable

Experiment and A/B Test Analysis

Check the design supports the causal claim before analysing it: ratio mismatch, randomisation, power, guardrails, and the honest rollout expectation.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Experiments are the one setting where causal claims are legitimate, and that licence is conditional. It comes from random assignment making the arms equivalent in expectation on everything, measured and unmeasured. Every serious failure in experiment analysis is a failure of one of the conditions behind that sentence: assignment did not really randomise, or randomised and then broke; the metric reported is the one that moved rather than the one specified; someone watched the result and stopped when it crossed; the test could never have detected the effect that matters; the arms interfered; a subgroup was mined until something appeared.

This skill runs the checks in the order that catches the most damaging failures first. Sample ratio mismatch comes before anything interesting, because it is cheap, decisive, and the strongest signal of a broken pipeline. Then randomisation verification and balance, clustering at the randomisation unit, the primary metric, guardrails, peeking, and the minimum detectable effect that makes a null interpretable. It handles novelty, interference, heterogeneous effects and the winner's curse, and it treats quasi-experimental designs as arguments whose identifying assumptions must be named and attacked.

It produces a validity block, an effect with its interval, and a rollout expectation that differs from the test result and says why.

Best used for

  • Deciding whether a completed A/B test supports the claim being made
  • Diagnosing a sample ratio mismatch before anything else is analysed
  • Interpreting a null result correctly instead of concluding no effect
  • Assessing whether a subgroup effect is real or the product of mining
  • Estimating an effect where randomisation was impossible, with assumptions checked
  • Explaining a gap between a test result and a rollout result
  • Setting the analysis rules for a test before it starts

Typical inputs

What you give it.

Unit of randomisation and the assignment mechanism, Pre-specified primary metric and decision rule, Unit counts per arm and the intended split, Outcome data at the unit of randomisation, Test window, deployment dates and any mid-flight changes, Pre-period outcome data for both arms (optional but decisive), Power calculation and its assumptions (optional), Pre-treatment covariates and guardrail definitions (optional), Rollout data after launch (optional)

Typical outputs

What you get back.

Validity block with sample ratio, balance, clustering and looks recorded, Power statement naming the smallest detectable effect against the material effect, Primary result with absolute and relative effect and confidence interval, Guardrail table with harm assessment, Secondary metrics with the count of metrics monitored, Subgroup table separating pre-specified from exploratory with interaction tests, Effect by exposure time and a persistence statement, Interference assessment naming the specific mechanism, Identifying assumption block for quasi-experimental designs, Rollout expectation stated as a range below the point estimate

Method coverage

What the skill works through.

  1. What licenses a causal claim, and what does not
  2. Sample ratio mismatch: the first check and the strongest signal
  3. Verifying that randomisation happened, and that it worked
  4. The unit of randomisation and the clustering it implies
  5. The pre-specified primary metric against the metric that moved
  6. Guardrail metrics and why they are analysed differently
  7. Peeking, optional stopping and inflated false positives
  8. Minimum detectable effect: could this test ever have answered the question
  9. Reporting the effect in absolute and relative terms
  10. Novelty, primacy and effects that change over exposure time
  11. Interference between arms and when unit-level randomisation breaks
  12. Heterogeneous effects without subgroup mining
  13. Reading a null result correctly
  14. Quasi-experimental designs and their identifying assumptions
  15. The gap between a test result and a rollout result

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

What is sample ratio mismatch and why does it matter so much?

It is a statistically significant deviation between the observed split of units across arms and the split you configured. It matters because the deviation is almost never random: it usually means units were lost, misassigned or double-counted by a mechanism that differs by arm, which destroys comparability in an unknown direction. Common causes are a redirect or load failure hitting one arm, uneven bot filtering, a logging bug, or units entering the dataset conditional on something the treatment affects. A test with an unexplained mismatch should be diagnosed, not analysed.

Do balance checks prove randomisation worked?

No, and the asymmetry is important. Balance checks can reveal that randomisation failed, but passing them does not prove it succeeded, because the whole value of randomisation is balancing the unmeasured variables you cannot check. Read the pattern rather than individual results: with twenty covariates at a 5% threshold you expect one flag, so a single imbalance is not evidence of failure. Several imbalances, or one large imbalance on a strong predictor of the outcome, is. The most powerful single check is the pre-period value of the outcome metric itself.

Why can I not stop a test as soon as it reaches significance?

Because repeatedly checking and stopping at the first crossing inflates the false positive rate substantially: with enough looks, a random walk crosses the threshold eventually. It also biases the effect estimate upward, because you stopped at a high point. If you need to monitor, use a design built for it, either group sequential boundaries with pre-specified interim looks or always-valid confidence sequences, and declare it in advance. Informal peeking followed by a nominal p-value is reporting a number that is not what it claims to be.

Does a non-significant A/B test mean the change does nothing?

No. It means the test could not distinguish the effect from zero, which is what a real effect looks like when the sample is too small. The honest statement combines three quantities: the point estimate, the confidence interval, and the minimum detectable effect at the sample achieved. If the interval excludes the effect size that would matter, that is genuinely informative. If it contains it, the test is inconclusive, not negative. The phrase "no difference" should not appear.

How should I handle subgroup results?

Distinguish pre-specified from exploratory, and test the interaction rather than the subgroups separately. Observing that an effect is significant among new users and not among returning users is not evidence that the two differ; that comparison is itself a test and must be run as one. Subgroup analysis also splits the sample, so a test adequately powered overall is usually underpowered for any subgroup, and the subgroup-level minimum detectable effect should be reported alongside the result. Exploratory subgroups are hypotheses for a confirmatory test, with the number of splits examined stated.

Why did our rollout not match the test result?

Several named mechanisms, usually more than one at once. The tested population was often more engaged than the rollout population. Novelty effects decay. The control group that existed during the test does not exist afterwards, so any benefit that came from taking share from control disappears. Capacity constraints that did not bind at test scale bind at full scale. And the winner's curse: a result selected as the best of several is biased upward. Expect the shipped effect below the tested one, state the range before launch, and measure the rollout so future tests can be calibrated.

What is interference and how do I know if it affects my test?

Interference is any situation where a unit's outcome depends on another unit's assignment, which breaks the core assumption of the analysis. It happens whenever arms compete for a finite shared resource (marketplace inventory, ad budget, staff time, stock), whenever units are socially connected, and whenever treated users influence untreated ones. Under interference the measured difference can be much larger or much smaller than the true effect. Assess the specific mechanism for your test rather than writing a generic sentence, and where it is plausible, randomise at a level that contains the interaction, such as geography or time period.

Can I make a causal claim without randomising?

Sometimes, conditionally, and the condition is an identifying assumption that must be named and checked. Difference-in-differences assumes parallel trends, testable with multiple pre-periods and untestable with one. Interrupted time series assumes nothing else changed at the discontinuity. Regression discontinuity assumes units cannot precisely manipulate their position around the threshold, checkable by looking for bunching. Matching and synthetic control assume no unobserved confounder remains, which is not testable. In every case, report the design, the assumption, the check and what would break it, in the same paragraph as the claim.

Should I randomise by user or by session?

Almost always by the coarser unit that matches how the treatment is experienced. Randomising sessions when a person has many sessions means they see an inconsistent experience, which contaminates the comparison. Randomising by user and then analysing sessions as if they were independent understates the standard error, sometimes severely, because observations within a user are correlated. Either aggregate to the randomisation unit or use cluster-robust methods, and report the intra-cluster correlation.

Can AI analyse experiments reliably?

It runs the checks consistently and without fatigue, which is exactly where human analysis tends to fail, since the validity checks are tedious and the effect estimate is the interesting part. The specific risks are producing an effect estimate before checking the sample ratio, using causal language without naming the design in the same sentence, reporting intervals computed as if clustered data were independent, presenting a post-hoc metric as the result, and treating a quasi-experimental estimate as though randomisation had occurred. Require the validity block first, the primary metric reported regardless of outcome, and the identifying assumption beside any non-randomised claim.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out