- Home
- Research skills
- Segment Comparison
Segment Comparison
Compare groups honestly: adequate bases, real tests, multiplicity controlled, composition confounds caught, and the sameness reported alongside the differences.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
A segment-by-measure table is the most over-read artefact in research. It is generated automatically, it contains hundreds of cells, and every cell invites a comparison nobody has tested. Four failures follow and they compound. Differences get declared by eye, so a five-point gap on bases of 90 and 400 becomes a strategic distinction. Six segments across forty measures produces hundreds of implicit tests, and at a 95% threshold a sizeable number of "significant" results are expected by chance with no way to tell which. An index of 180 looks decisive and can rest on eleven respondents. And two segments differ on digital channel use because one skews fifteen years younger, not because of anything the segmentation defines.
This skill runs a difference through five gates in order: base, independence, reliability, composition, materiality. It labels the comparisons that are circular, because a segmentation compared on the variables used to build it proves nothing. It separates a difference in level, where groups want the same things at different intensities, from a difference in structure, where priorities genuinely differ. And it requires a shared-ground section, because comparison tables structurally overstate difference and most groups have far more in common than the table can show.
You get a base grid, tested differences with composition checks, a corrected index table, an explicit exploratory set, and a record from which any reported difference can be reconstructed.
Best used for
- Profiling segments on measures used and not used to build them
- Comparing screener-defined, database-defined or naturally occurring groups
- Testing whether a striking segment difference is really an age or market effect
- Handling the multiple-comparison problem in a large segment-by-measure grid
- Correcting index scores read as group size
- Distinguishing a difference in level from a difference in priority structure
- Auditing a supplied comparison table against a five-gate standard
Typical inputs
What you give it.
Group definitions and how the groups were created, Base size and base description for every group on every measure, The measures with their wording, scales and fieldwork conditions, The basis variable list, for segmentation-derived groups, The decision the comparison informs, Demographic and structural profile of each group (optional), Weighting variables and effective base per group (optional), A pre-registered analysis plan or hypothesis list (optional), Prior wave comparisons (optional), Total-sample figures for index calculation (optional)
Typical outputs
What you get back.
Comparison design statement covering planned versus exploratory and the multiplicity treatment, Base grid flagging every group below 100 and below 30 on every measure, Circular comparison block, labelled as differences by construction, Difference table with tests, thresholds, composition checks, level or structure, and materiality, Index table showing index, underlying percentage, base and group size together, Shared-ground section listing decision-relevant measures on which groups do not differ, Real but immaterial differences, listed separately, Exploratory findings labelled as hypotheses with the number of tests stated, Statement of what could not be compared and why
Method coverage
What the skill works through.
- What makes a comparison meaningful rather than merely available
- The base grid, and why the weakest group governs the row
- Circular comparisons: profiling segments on the variables that built them
- Testing rather than eyeballing, and what comparative language is permitted without a test
- The multiple-comparison problem in a segment-by-measure grid
- Three legitimate responses to multiplicity: restrict, adjust, disclose
- The composition confound and how to check for it
- When a difference disappears under a control, and why that is still a finding
- Difference in level versus difference in structure
- Index scores: what they hide and how to present them safely
- The high index on a low base
- Statistically real and practically irrelevant
- Reporting what the groups share, not only what divides them
- Auditing a supplied comparison table
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
How big does a base need to be to compare two groups?
Below 30 in either group, no percentages are reported and no comparison is made. Between 30 and 99, the comparison is directional only and a difference has to be large to be reliable. The rule that gets forgotten is that the weakest group governs the whole row: a group of 62 compared with a group of 900 is a weak comparison, not a strong one with an asterisk, and a strong base on one side does not compensate. Build a base grid before any analysis, because bases move with routing and a group's base on one question is not its base on another.
What is the multiple-comparison problem in segment analysis?
Six segments compared pairwise across forty measures generates hundreds of tests. At a 95% threshold, roughly one in twenty of the genuinely null comparisons will pass by chance, so a large grid can be expected to produce dozens of spurious results that look exactly like real ones. There are three legitimate responses. Restrict: decide which comparisons matter before looking, and test only those. Adjust: apply a correction across the family of tests, accepting reduced sensitivity. Disclose: report how many comparisons were made and how many are expected by chance, and label the set as hypothesis-generating. Running everything at 95% and reporting the winners is not one of the three.
What does over-indexing actually mean?
That a group's figure on a measure is higher than the total, expressed as a ratio. It is useful for spotting concentration in a large table and it misleads in three ways. It hides the absolute value: an index of 200 where the total is 3% means 6%, which is almost nobody. It hides the base: small bases produce extreme indices as a mechanical consequence of instability. And it hides the group's size: a strongly over-indexing segment of 4% of the market contains fewer people than a mildly under-indexing segment of 40%. Always show the index with its underlying percentage, its base and the group's size, never index on a base below 100, and never rank opportunities by index alone.
My segment looks completely different on digital behaviour. Is that real?
Check the composition first. If the segment is fifteen years younger than the rest of the sample, it will differ on digital channel use, media consumption and price sensitivity, and none of those differences is caused by whatever the segmentation was built on. Compare within age bands: does the gap survive when you compare like with like? If it disappears, the honest finding is that the groups differ in composition and the behaviour follows from that. This is not a downgrade; it is a more useful finding, because it points at targeting by age or by observed behaviour rather than by segment.
What is the difference between a level difference and a structure difference?
A level difference is when two groups give different absolute answers but agree on rank order and relative gaps: they want the same things in the same order, one group is simply more positive, more engaged or more generous with the scale. A structure difference is when the ordering itself changes: one group ranks price first and convenience fourth, the other the reverse. Level differences are far more common and are routinely reported as structural, which produces differentiated strategies for groups that actually want the same thing. Compare rank orders, not only absolute values.
Is it valid to compare segments on the questions used to create them?
It is valid to show, and worthless as evidence. The segments differ on those variables because they were constructed to differ on them. Showing the block is useful because it describes what defines the groups, but it must be labelled as differences by construction. The uncomfortable corollary is that a segmentation whose only large differences are circular has not been shown to describe anything real, which is a finding about the segmentation rather than about the audience.
A difference is statistically significant but tiny. Should I report it?
Report it in a clearly separated list marked as real but immaterial, so the reader does not have to make that judgement across forty rows. Large samples make small differences reliable, and reporting every reliable difference produces a document nobody can prioritise. The discipline is to set a materiality threshold before you see the numbers, by asking how big a difference would change the decision. Set afterwards, the threshold becomes whatever the data happened to produce.
Why should I report what segments have in common?
Because a comparison table can only show you differences, so it systematically overstates how different the groups are. In most studies, the list of decision-relevant measures on which groups do not differ is longer than the list on which they do, and it is directly actionable: it tells the business which parts of the product, proposition and service can be common. A difference-only output produces a management team that believes it has five markets, and a cost base to match.
How do I audit a comparison table someone else produced?
Run the five gates against it and produce a marked-up version showing which differences survive each one. The base grid alone usually removes a substantial share of the claimed differences. The circularity check removes more. The composition check changes the interpretation of several survivors rather than removing them. This is more persuasive than rebuilding the analysis, and it gives the client a standard they can apply to the next table they receive.
Why do market comparisons produce so many differences?
Partly because cultures use rating scales differently, which is large, systematic and easily mistaken for a difference in attitude. If an entire battery shifts in one direction for one market, suspect response style before substance. Standardising within respondent removes most of it and some real variance with it, so report both versions. Comparability of wording, translation, mode and fieldwork period also has to be established before any market comparison means anything, and the cultural reading of what survives belongs to a human with market knowledge.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
