09.01Segmentation and Audience UnderstandingAvailable

Audience Segmentation

Build segments that are real, stable and actionable, starting with the question most segmentations skip: is this market segmentable at all?

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

A clustering algorithm always returns clusters. Give it random numbers and it will partition them, report separation statistics, and hand back groups you can name, colour and put on a wall. Nothing in the standard workflow asks whether the partition corresponds to anything in the world, which is why so many organisations spend years building products and campaigns for segments that do not exist.

This skill puts the prior question first. Before any solution is run, it assesses whether the population is genuinely heterogeneous on the dimensions the decision depends on, and it treats "this market is a spectrum, not a set of groups" as a legitimate and valuable result. Where a segmentation is warranted, it chooses basis variables from the decision rather than from whatever was in the questionnaire, keeps demographics as descriptors where they belong, runs many solutions instead of one, and holds a stability standard that most solutions quietly fail.

You get a segmentability assessment, a solution matrix with the rejected solutions and why they failed, three stability tests with their results, segment sizes in value as well as in people, a profile built on variables nobody used to construct the segments, and a typing tool that publishes its own error rate.

Best used for

  • Establishing whether a category is segmentable before money is committed
  • Building needs-based, attitudinal, behavioural or occasion-based segmentations
  • Deciding how many segments the data actually supports
  • Auditing an existing segmentation that nobody uses or nobody believes
  • Testing a segmentation for stability rather than accepting separation statistics
  • Building and validating a typing tool, with its misclassification rate
  • Sizing segments in value rather than in respondents
  • Reporting honestly that a market is a spectrum rather than a set of groups

Typical inputs

What you give it.

The decision the segmentation must serve, and who owns it, Respondent-level dataset with candidate basis variables measured on the whole sample, Questionnaire as fielded, with scale direction and routing, Data dictionary including derived-variable construction rules, Sample and fieldwork documentation, Linked behavioural or transactional data (optional), Qualitative needs evidence from the same population (optional), Prior segmentation and its typing tool (optional), Hold-out sample or second wave (optional), Customer database schema and targeting capability (optional)

Typical outputs

What you get back.

Segmentability assessment stating whether the market divides usefully at all, Basis and descriptor register with a justification per variable, Solution matrix across multiple algorithms and multiple numbers of segments, Stability report covering split-half, alternative algorithm and alternative variable subset, Chosen solution with the rejected solutions and the gate each failed, Segment descriptions with sizes in people and in category value, External profile on variables not used to build the solution, with tested differences, Typing tool with confusion matrix and per-segment accuracy, Action table mapping each segment to what the business would do differently, Limitations, currency statement and re-validation trigger

Method coverage

What the skill works through.

  1. The prior question: is this market segmentable at all?
  2. Why clustering always returns clusters, and what that means for your evidence
  3. Choosing basis variables from the decision the segmentation must serve
  4. Why demographics are descriptors and almost never bases
  5. Scale-use bias, and the segments it manufactures
  6. Running many solutions instead of one
  7. The four gates: separation, stability, size, differentiability of action
  8. Stability testing: split-half, alternative algorithm, alternative variable set
  9. The reportability standard for an unstable solution
  10. Sizing segments in value, and what counts as too small to serve
  11. Profiling on variables you did not use to build the segments
  12. Naming segments without asserting more than you measured
  13. Typing tools and the misclassification rate you have to publish
  14. Segment decay, and when to re-validate
  15. Where a researcher's judgement is required

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How do I know if my market is actually segmentable?

Run four checks before building anything. Look at the distribution shape of every candidate basis variable: unimodal, symmetric distributions across the whole battery indicate a population that varies by degree rather than by kind. Look at the correlation structure: if the battery collapses to one or two dimensions on which everyone sits somewhere along a continuum, you have a spectrum, not segments. Look at the qualitative evidence: does it describe different kinds of people with different logics, or the same person at different levels of engagement? And look at behavioural dispersion: do people differ in what they do, or only in how much? Where three of the four point to homogeneity, the honest deliverable is that finding, not a forced solution.

Does statistical separation prove that my segments are real?

No, and this is the central misunderstanding in the field. Separation is produced by the method regardless of what the data contains, so it cannot be evidence about the data. Two things carry the reality claim instead: stability, meaning the same structure appears under different sample splits, different algorithms and different subsets of the basis variables; and external profiling, meaning the segments differ on variables that were not used to build them. A solution with modest separation that reappears under every method is worth more than a sharp one that dissolves when the algorithm changes.

How many segments should a segmentation have?

As many as the data supports, which is frequently fewer than the number requested. Run a range wider than the client's expectation, score every solution on separation, stability, size and whether the business would act differently towards each group, and let any one of those four disqualify a solution. If a six-segment solution has two groups that swap members freely under a different algorithm and a four-segment solution reproduces across splits, report four, and report what happens at six and why it was rejected.

Should demographics be used to build segments?

Almost never. Age, gender, income, region and firmographics are cheap to collect and easy to target, but they are rarely the reason anyone behaves as they do. Build on needs, attitudes, behaviour, occasions or value, then profile demographically afterwards so the segments can be found and reached. Reversing this order is the most common structural error in segmentation practice, and it produces segments that are essentially age bands with attitude labels attached. Demographics earn a place as a basis only where they are the mechanism itself, such as life stage in a category defined by life stage.

How do I test whether a segmentation is stable?

Three ways. Split the sample randomly, build the solution independently on each half, apply each half's rule to the other and measure agreement. Build the same number of segments with a different algorithm and cross-tabulate the two allocations. Rebuild on a random subset of the basis variables and compare. What you are looking for is consistency of structure: the same number of recognisable groups with the same defining differences, and individuals mostly landing in corresponding groups. If segment shape changes when the method changes, the structure belongs to the method, not the market, and it is not reportable.

What is a typing tool and why does its accuracy matter?

A typing tool is a short set of questions plus a rule that allocates a new individual to a segment, so segments can be applied to a database, a media plan or a later wave. Every typing tool misclassifies some people, and the rate is a property of the tool rather than a defect to be hidden. Report the overall hit rate and the rate per segment, because a tool that is 80% accurate overall can be 50% accurate on the smallest and most strategically interesting group. Once allocations enter a customer database they are treated as facts about the person, and the misallocated cases become indistinguishable from the rest.

What makes a segmentation actionable?

Two things, both external to the statistics. The segments must be reachable, meaning you can identify who belongs to which group in the real world through a database, a channel or a media plan. And the business must genuinely do something different for each of them: a different product, message, price, channel or service level. Write that action table before you see the data. If two segments have the same action column, they are one segment for this decision, however well they separate.

Why do my segments not feel real to stakeholders?

Usually one of three reasons. They were built on a homogeneous market, so the differences are arbitrary and nobody can describe a segment without using the word "more". They differ only on the variables used to construct them, so they are a restatement of the battery rather than a discovery about people. Or the descriptions assert character that was never measured, so they read as fiction to anyone who knows the customers. The diagnostic is to profile the segments on data nobody used to build them and see whether the differences survive.

How often should a segmentation be refreshed?

Attach a re-validation trigger rather than a schedule: a category shift, a competitive entry, a change in the buying process, or a fixed interval of two to three years, whichever comes first. Re-validate by applying the existing typing tool to a fresh sample and independently rebuilding a solution on the same basis variables. Either the structure holds and only sizes have moved, or membership has migrated and the tool needs re-issuing, or the structure no longer appears at all, at which point replacement becomes a business decision about switching costs rather than a research one.

Can AI build a segmentation on its own?

It can run the solutions faster and more consistently than a person, and it will reliably produce the number of segments requested, which is exactly the problem. The judgements that make a segmentation defensible are not computational: whether the differences are commercially material, whether the naming survives cultural reading in each market, whether the organisation can act on the structure, and whether an unstable solution should be reported as a failure. Those return to a named researcher, and the output should mark them rather than resolve them.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out