09.03Segmentation and Audience UnderstandingAvailable

Behavioural Profiling

Build a picture of what people actually do: observed separated from claimed, repertoires instead of loyalty, occasions instead of averages, light buyers included.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Most descriptions of customer behaviour are descriptions of what customers said about their behaviour. People do not count their own purchases; they estimate, and the estimate is wrong in known directions. Regular behaviours are over-reported, irregular ones under-reported, approved ones inflated, and older events pulled forward into the recall window. The longer the recall period, the worse all four get. A profile built on those answers is not noisy, it is biased, and the bias is largest exactly where the business is most interested.

This skill ranks the available evidence on reliability before analysing any of it, and keeps the observed and claimed labels attached to every figure through to the last slide, which is where they normally get lost. It replaces self-reported loyalty with repertoire size and share of category, because in most repeat-purchase categories people use several brands and exclusive buyers are a small minority. It tests whether the person or the occasion is the right unit before profiling anything. And it puts the full buyer distribution in front of any heavy-user analysis, because the light-buyer majority usually carries more volume than the heavy minority and is routinely designed out of the strategy.

You get a sourced behavioural profile, a measured claimed-versus-observed gap reported as a finding rather than reconciled away, repertoire and share-of-category figures, and an explicit list of what the behavioural record cannot see.

Best used for

  • Building usage and behaviour profiles that separate observed from claimed
  • Reconciling survey claims with transactional or telemetry records
  • Replacing self-reported loyalty with repertoire and share of category
  • Establishing whether the person or the occasion is the right unit of analysis
  • Correcting strategies built on heavy users without the light-buyer volume share
  • Deriving behavioural groups from behaviour and testing their persistence
  • Distinguishing a changed person from a changed situation between waves

Typical inputs

What you give it.

At least one behavioural evidence source with its collection method documented, Population definition and coverage boundary for each source, The period each source covers, and whether it was ordinary, The category and event definitions in use, A second independent behavioural source for claimed versus observed comparison (optional), Occasion-level or in-the-moment data such as diary or experience sampling (optional), Competitive or category-wide behavioural data for share of category (optional), Longitudinal or panel data on the same individuals (optional), Attitudinal measures on the same respondents (optional), Household or account structure (optional)

Typical outputs

What you get back.

Source inventory with reliability rank, coverage boundary, period and unit, Definitions block that would reproduce every figure from raw records, Behavioural profile table marking each figure observed or claimed, with base and period, Buyer distribution showing buyers and volume share by frequency band, Repertoire size, share of category requirements and penetration, Occasion profile where occasion-level data exists, Measured claimed-versus-observed gap, reported as a finding, Behavioural groups with their persistence across periods, Attitude associations with direction explicitly unestablished, Blind-spot list and statement of what could not be established

Method coverage

What the skill works through.

  1. Ranking behavioural evidence: observed, in-the-moment, short recall, long recall
  2. Why claimed frequency is an estimate and not a count
  3. The four predictable errors: rate heuristics, telescoping, social desirability, window length
  4. Measuring the gap between claimed and observed instead of choosing a side
  5. Person or occasion: testing where the variance actually lives
  6. Repertoire, share of category requirements and penetration
  7. Why self-reported loyalty measures self-image
  8. The buyer distribution, and the light-buyer volume share
  9. The heavy-user fallacy and what it costs
  10. Deriving behavioural groups from behaviour, and checking they persist
  11. Changed people or changed circumstances: what a cross-sectional design cannot tell you
  12. Linking behaviour to attitude without asserting a cause
  13. Coverage boundaries and the blind spots of every behavioural record
  14. Reporting so that observed and claimed survive into the summary

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How accurate is claimed purchase frequency in surveys?

Less accurate than it looks, and biased rather than merely imprecise. People answer by applying a rate heuristic ("about twice a month") or by recalling instances and extrapolating, so regular behaviours are over-reported, irregular ones are under-reported, socially approved behaviours are inflated, and events from before the window are pulled into it. All four effects worsen as the recall window lengthens, so a twelve-month frequency claim carries much less information than a seven-day one. The most important operational rule is never to multiply a claimed frequency into an annual or market volume, because that compounds the bias into a number that then enters a business case as a fact.

What do I do when survey answers and transactional data disagree?

Measure the gap rather than choosing a side. Cross-tabulate claimed against observed where the data links at individual level, and report the size, direction and pattern of the difference. The gap is usually patterned: heavier users over-estimate more, light users less, so the two errors together compress the true distribution towards the middle. Treat that as a finding, because it tells you how the audience understands its own behaviour, which is directly relevant to how they will respond to any communication about it. Do not average the two sources, and do not silently prefer the survey because it has more variables attached to it.

Should I measure brand loyalty?

Usually not in the form it is normally asked. In most repeat-purchase categories buyers hold a repertoire of several brands and allocate volume across it, and exclusive buyers are a minority even among a brand's own customers. A self-reported "main brand" question therefore measures self-image rather than behaviour. Report repertoire size, share of category requirements and penetration instead. Penetration and share of requirements separate two different growth routes, more buyers or more share from existing buyers, which a single loyalty number hides.

Why is profiling the heavy user a mistake?

Because heavy users are unrepresentative by construction, and in most measured categories the large majority of buyers buy infrequently and collectively account for a substantial share of volume, often more than the heavy minority. A proposition designed around heavy users addresses the people least in need of persuading, and the arithmetic of growth usually favours reaching more light buyers. Report the full buyer distribution, showing share of buyers and share of volume by frequency band, before any heavy-user content appears. Heavy-user analysis is legitimate; it is a description of a minority and should be framed as one.

Should I profile people or occasions?

Test it rather than assuming. Where occasion-level data exists, compare within-person variation with between-person variation on the key behaviour. In many categories the same individual behaves quite differently at different times, and the within-person variation is larger, which means a person-level profile produces an average that fits no actual situation and a segmentation that will not stabilise. Where that is the case, profile occasions and describe people by their occasion repertoire. Where occasion data does not exist, say that the unit could not be tested.

Does an attitude-behaviour correlation show that attitude drives behaviour?

No, and in behavioural research the reverse direction is not a technicality. People report warmer attitudes towards brands they already use, so a correlation between brand attitude and purchase frequency is at least as consistent with usage producing attitude as the other way round. Report the association with its base, state that the direction is not established by a cross-sectional design, and name the design that would settle it if the answer matters to the decision.

Behaviour changed between waves. Did people change?

You cannot tell from cross-sectional data. Three explanations compete: the same people changed, different people are in the sample, or the same people are in different circumstances such as a price change, a supply problem or a seasonal effect. Only panel data on the same individuals separates the first from the second. Where you have it, decompose the change into buyers gained and lost versus rate changes among retained buyers; that decomposition is usually more actionable than the headline movement.

What are the blind spots of transactional data?

Coverage, not accuracy. A single retailer's file describes purchases at that retailer; an app's telemetry describes use of that app; a card record misses cash. It also usually resolves to a household or an account rather than a person, which inflates every per-person frequency, and it carries no context about why anything happened. State the coverage boundary on every figure derived from it, and never infer that a customer buys nowhere else from the absence of competitor records.

How should behavioural groups be derived?

From recorded behaviour: frequency, recency, repertoire, share of category, occasion mix, channel mix. Not from a self-classification item, which measures self-image. Then check persistence: someone in the heaviest band this quarter is frequently in a middle band next quarter, purely through regression to the mean. A behavioural group that does not persist across periods describes a period rather than a group of people, and it will not support a targeting strategy.

Can I build a behavioural profile from survey data alone?

Yes, provided it is honest about what it is. Label every figure as claimed, state the likely direction of bias for each measure, use the shortest recall window available, report distributions rather than means on skewed behaviours, and calculate no derived volumes from it. Confidence caps at moderate for relative statements and low for absolute levels. Then name the observed source that would fix it, because that sentence is often what gets the data acquired for the next study.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out