05.03Quantitative AnalysisAvailable

Cross-Tabulation

Design a banner that answers the analysis plan, base every cell, percentage in the right direction, and stop a wide banner manufacturing false findings.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

The cross-tab is the workhorse of quantitative research and the place where most false findings are made. Three failures do nearly all the damage. The banner nobody designed, where every demographic becomes a column because the software offered it. The direction error, where a table percentaged across the row is described as though it were percentaged down the column, which does not blur a finding, it inverts it. And the harvest, where an analyst scans a banner for flagged cells and writes up whatever came up.

This skill supplies the design discipline before the table is run and the reading discipline after it. It derives the banner from the objectives and records an expectation for each column before the tables exist. It defines every banner point in code and tests whether columns overlap. It puts unweighted and effective bases in the table rather than a footnote, and reports a cell too small to publish as too small rather than leaving it out. It teaches reading across a row before reacting to a cell, checking whether a difference is a composition effect, counting the comparisons a banner generates, and separating a difference you can see from one you can claim.

Best used for

  • Specifying a tabulation job before it runs
  • Deciding which banner points belong in the analysis
  • Answering whether a result varies by segment, tenure, usage or region
  • Reading a banner somebody else produced and deciding what is reportable
  • Auditing a deck built from cross-tabs for direction and base errors
  • Preparing the comparison set for significance testing
  • Diagnosing whether a group difference is a composition effect

Typical inputs

What you give it.

Prepared respondent-level dataset, Base register or routing and filter logic from descriptive analysis, Questionnaire as fielded, with scale direction and response lists, Research objectives or analysis plan, Definitions in code for every proposed banner point, Weighting scheme and effective bases per column (optional), Pre-specified comparisons from the analysis plan (optional), Previous wave banner specification (optional), Segment membership variables (optional)

Typical outputs

What you get back.

Banner specification with objective and expectation per column, Cross-tabulation tables titled with base, direction of percentaging and weighting status, Base row carrying unweighted n and effective base per column, Row reading notes describing the shape of each row taken forward, Comparison count and multiplicity statement, Three labelled difference lists (pre-specified confirmed, pre-specified not found, exploratory), Suppression register of every cell held back with its base and reason, Conventions block for handover to testing and reporting

Method coverage

What the skill works through.

  1. What a cross-tab is for, and what it cannot do
  2. Designing a banner from the analysis plan rather than the variable list
  3. Writing the expectation before the table exists
  4. Defining banner points in code: exclusivity, exhaustivity and overlap
  5. Column, row and total percentages, and how getting the direction wrong inverts a finding
  6. Base sizes in every cell, and effective bases on weighted data
  7. Reporting a cell that is too small to report
  8. Reading across the row before reacting to a cell
  9. Gradients versus spikes
  10. When a group difference is really a composition effect
  11. How many comparisons a banner generates, and how many will flag by chance
  12. Separating a difference you can see from a difference you can claim
  13. What must travel with a set of tables

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How do I decide what goes in a cross-tab banner?

Start from the questions the study exists to answer and write, against each, the comparison that would answer it. A banner point earns its column by appearing in that list. Adding every available demographic looks thorough and is the opposite: it dilutes the analysis, multiplies the comparisons and guarantees false positives. Keep the excluded variables in a second-tier list to be run only if a first-tier result needs explaining.

What is the difference between column percentages and row percentages?

Column percentages answer "of the people in this group, what proportion gave this answer" and each column sums to 100. Row percentages answer "of the people who gave this answer, what proportion are in this group" and each row sums to 100. A product bought by 4% of under-35s and 1% of over-55s is four times more popular among the young, and most of its buyers may still be older, because there are more older people. Both statements are true and they come from different directions of the same table. Reporting one while describing the other is the commonest false statement in commercial research.

What base size is too small to report in a cross-tab?

Below 30 in a cell, no percentages at all: report counts, or report the cell as too small. Between 30 and 99, show the base, flag it and read directionally only. On weighted data the threshold applies to the effective base, not the nominal one, so a column of 120 with a design effect of 1.7 behaves like a column of 70. The important rule is what happens below the threshold: a cell too small to report is reported as too small, with its base, rather than left out. A blank tells the reader nothing was found there, which is a claim, and usually a false one.

Why are there so many significant differences in my cross-tab?

Because a banner runs an enormous number of comparisons. Thirteen columns produce 78 pairwise comparisons per row; thirty columns produce 435. A thirty-column banner across forty questions is well over a thousand comparisons even on a disciplined reading, and at a 5% threshold roughly one in twenty flags with nothing true anywhere. The false positives are arithmetic, not bad luck. Count the comparisons before running them, report the count alongside the findings, and label anything that was not predicted in advance as exploratory.

How do I read a cross-tab properly?

Read the base row first, then read across the row rather than down onto whichever number is furthest from the total. Ask what shape the row has. A steady gradient across ordered columns is credible, because noise does not usually arrange itself in order. A single column standing away from a flat row is a hypothesis, and it is very often the cell with the smallest base, since small bases produce large swings. Only after the shape of the row is described should any individual cell be singled out.

Can a subgroup difference in a cross-tab be caused by something else?

Routinely. Banner points are correlated with each other, so a regional difference may be an age difference wearing a regional label. The diagnostic is to nest one banner point inside the other and see whether the gap survives. In the extreme case a relationship can hold inside every subgroup and reverse in the combined table, because the subgroups differ in size and in base rate. A cross-tab cannot separate correlated variables, so either run the nested check, or say the difference is confounded and move to a model.

Can I compare two columns in my banner directly?

Only if a respondent cannot appear in both. Age bands are fine. "Users of product A" and "users of product B" usually are not, because people use both, and comparing overlapping groups with a standard two-sample test inflates significance. Test the exclusivity in the data rather than assuming it from the labels, and mark any overlapping column in the header so nobody compares it later.

Does a cross-tab tell me whether a difference is real?

No. It shows you the difference. Whether the difference is larger than sampling variation is a separate question requiring a test, and that boundary is where cross-tabulation ends. Describing a gap arithmetically with both bases shown is legitimate. Words like "notably", "clearly" and "significantly" are claims about reliability and need a test behind them, whatever flags the tabulation software has printed.

Can AI produce and read cross-tabs reliably?

It applies conventions consistently, which is a real advantage on a large tabulation. The risks are specific: building a banner from whatever variables exist rather than from the objectives, narrating every gap in the table as though each were a finding, describing a percentage in language that matches the other direction of percentaging, and inferring what a banner point means from its label. Used properly, AI should be required to state the direction of percentaging and the base with every figure, report the count of comparisons its findings were selected from, and label anything unpredicted as a hypothesis.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out