05.06Quantitative AnalysisAvailable

Correlation, Regression and Causal Claim Control

Read relationships properly, recognise the structures that make an association misleading, and know exactly which designs license a causal claim.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Almost every research finding that gets acted on is acted on causally. The report says two things are associated; the meeting decides to change one in order to move the other. That step is taken silently, all the time, and usually nothing in the analysis licensed it. Reciting that correlation is not causation helps with none of the questions that matter: which structures produce a misleading association, how to recognise one in your own data, what "controlling for" achieves and cannot, why adding a control sometimes makes a result worse, and which designs license the claim.

This skill covers both halves. On relationships: what a correlation coefficient conceals, how range restriction and outliers distort it, the non-linearity it cannot show, and the fact that very different scatter patterns produce the same number. On regression: what a coefficient means, what conditioning achieves, how to read fit without over-reading it, and which assumptions matter.

Then the firewall: the five confounding structures explained so an analyst can recognise them in their own data; a language pass over a finished draft, headings and executive summary included; the designs that license causation and the assumptions each rests on; and how to answer a stakeholder who wants the claim anyway.

Best used for

  • Assessing a relationship between two measures and reading what the coefficient hides
  • Estimating an association while holding other variables constant
  • Diagnosing why a coefficient changed sign when a variable was added
  • Recognising confounding, mediation, collider and selection structures in real data
  • Running a causal language pass over a draft report or deck
  • Advising on the design that would license a causal claim
  • Responding to a stakeholder who wants an attribution the data cannot support

Typical inputs

What you give it.

The relationship question, with outcome and predictor of interest named, Respondent-level or record-level data on the same units and base, Measurement definition of every variable, including derivation rules, The design that produced the data and when each variable was measured, Analysis base after missing-data treatment, A stated prior structure of what is believed to cause what (optional), A source of exogenous variation such as a staged rollout or threshold (optional), Behavioural or transactional data alongside self-report (optional), A comparison group unaffected by the intervention (optional)

Typical outputs

What you get back.

Question block separating the relationship question from the causal question, Design statement naming what class of claim is available, Relationship description with shape, coefficient, range and influential cases, Model table with coefficients, intervals and the conditioning set in words, Confounding assessment covering common cause, mediator, collider, selection and reverse causation, Causal position naming the licensing design or stating that none exists, Language pass record over headings, chart titles and the executive summary, Explicit statement of the assumption made if the organisation acts anyway

Method coverage

What the skill works through.

  1. Why the correlation-is-not-causation slogan does not help
  2. Reading the scatter before quoting the coefficient
  3. Range restriction, outliers and non-linearity
  4. Why several different relationships produce the same coefficient
  5. Regression at a working level: what a coefficient actually means
  6. What "controlling for" achieves, and its four permanent limits
  7. Why adding a control variable can make an estimate worse
  8. Reading model fit without over-reading it
  9. The five confounding structures: common cause, mediator, collider, selection, reverse causation
  10. Recognising each one in your own data
  11. The designs that license a causal claim
  12. Randomised experiments, valid quasi-experiments, longitudinal designs
  13. The language pass over a finished draft
  14. Answering a stakeholder who wants the causal claim anyway

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

What does "controlling for" a variable actually do?

It removes the part of the association that runs through that variable, as measured. Four limits are permanent. It only removes what was measured, so an unmeasured confounder is untouched. It only removes it as well as the measure captures it, so a noisy measure leaves most of the confounding in place while creating the impression it has been handled. It can make things worse if the wrong variable is chosen. And a model with more controls is not more rigorous, it is a different model answering a different question.

Why did my coefficient change sign when I added a variable?

Usually one of three things. The added variable is a confounder and the original estimate was picking up its influence, in which case the new estimate is better. It is a mediator sitting on the path between your predictor and the outcome, in which case you have removed part of the effect you were trying to measure and the new estimate is worse. Or it is highly correlated with your predictor, in which case neither coefficient is stable and the sign is close to arbitrary. Deciding which requires a view of what causes what, not a diagnostic.

What is collider bias?

When two variables both influence a third, conditioning on that third variable creates an association between them that does not otherwise exist. It is the structure practitioners find least intuitive and it appears constantly, because colliders are frequently the variables that define a sample. If two independent qualities both raise the chance of appearing in your dataset, they will appear negatively associated inside it. The test is to ask whether your control variable, or the way units entered the sample, is a consequence of both variables of interest.

How do I know if a relationship is caused by selection?

Describe in words exactly how a unit came to be in the dataset. Current customers, active users, survey completers, people who finished a journey, firms that survived. If the entry condition is related to both variables you are studying, the association inside the sample can differ in size and even in sign from the association in the population. This is not an edge case; most commercial datasets are selected in some way.

Which designs allow a causal claim?

Three families. A randomised experiment, where assignment is random so the groups are alike in expectation on everything measured and unmeasured. A valid quasi-experiment, where something outside the units' control created variation that behaves as though random: a staged rollout, an eligibility threshold, a policy change affecting one group only. And a longitudinal design where the predictor was measured before the outcome, the baseline outcome is controlled, and the plausible confounders were measured. Each rests on assumptions that must be stated and evidenced rather than asserted, and the licensing design is named in the same paragraph as any causal claim.

Does a strong correlation make causation more likely?

No. The strength of an association tells you about precision and magnitude; it tells you nothing about direction or about whether a third variable produced both. A very strong association from a cross-sectional design licenses no causal claim, and a modest one from a randomised design licenses a claim about its measured effect. Practitioners consistently reverse this, treating a large or highly significant result as though its size were evidence about its cause.

How do I stop causal language creeping into a report?

Run a deliberate pass over the finished draft, including headings, chart titles, bullet fragments and the executive summary, not only the body. Mark every instance of "drives", "leads to", "results in", "the impact of", "improving X will", and the bare "so" or "therefore" joining two findings. Then catch the constructions with no verb: a heading of the form "why customers churn" over correlational evidence, a chart titled "what improves retention", an arrow in a diagram, a recommendation whose logic only works if the association is causal. The executive summary fails most often, because it is written last, compressed hardest and read most.

What do I say to a stakeholder who wants the causal claim?

Four moves in order. State what the evidence does support, precisely and without hedging, so the strength of what you have is not lost. Name the specific rival explanation rather than the general principle, because "the customers who use this feature were already going to renew, and you can see it in their tenure profile" persuades where a slogan does not. Name the design that would settle it, with a rough sense of cost and timing. Then offer the strongest honest formulation, including, where the decision must be made anyway, an explicit statement of what the organisation is assuming if it acts as though the relationship were causal. Making the assumption visible is usually the most useful thing you can produce.

Is a high R-squared a good sign?

It means the model predicts well within this sample, and nothing more. It is not the share of the outcome that is caused by the predictors, and it is not evidence that the model is correctly specified. It rises whenever a variable is added. A high value with a wrong specification is entirely possible, and a low value can accompany a correctly specified and important relationship, particularly where individual behaviour is inherently variable.

Can AI be trusted with causal reasoning in research?

It is good at applying the checks consistently and at catching causal constructions in a long document, which is genuinely useful. The risks are specific: it writes fluent explanatory narratives that read as established mechanism when they are hypotheses, it lets causal verbs survive in headings and summaries after removing them from the body, and it hedges everything uniformly so a reader cannot tell a strong finding from a weak one. Used properly, AI should be required to assess all five confounding structures against the specific dataset, name the licensing design beside any causal claim, and write the headline so that it survives being quoted without its caveats.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out