02.07Instrument DesignAvailable

Scale and Measurement Selection

Decide how each construct is measured: scale type, points, labels, midpoints, single or multi-item, and the statistic it will actually produce.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Scale decisions are usually made by habit and discovered at analysis. Someone reaches for a five-point agreement battery because the last study had one, and by the time anyone examines it the data has three properties nobody chose: a heavy acquiescence component, a distribution with most respondents in two points, and a mean that cannot be compared with anything.

Underneath the habit sit real decisions that were never made. Whether the thing being measured is one construct or several. Whether the scale runs from nothing to a lot or from one pole to its opposite. How many points, which is genuinely contested. Whether there is a real neutral position or the midpoint is a parking space. Whether an agreement frame is being used because agreement is the construct or because it was easy to write. And what statistic will be reported, which should have driven everything else and is almost always decided last.

This skill covers all of those, plus frequency scales and vague quantifiers, behavioural versus attitudinal measurement, when to adopt a validated battery instead of writing your own, the cost of changing a scale mid-tracker and how to bridge it, and the specific dangers of net and difference scores.

Best used for

  • Choosing response frames for a new questionnaire
  • Fixing an instrument with inconsistent scales and label sets
  • Deciding whether one item can carry a construct
  • Settling point count, labelling and midpoint arguments with reasons
  • Deciding whether to adopt an existing validated battery
  • Changing a scale in a tracker and bridging the break
  • Assessing whether a proposed net or difference score is defensible

Typical inputs

What you give it.

Construct definitions naming object, aspect and period, The planned analysis and the statistic each measure will produce, Mode and device profile, Previous wave or comparable instrument, where a trend exists (optional), External benchmarks the study must match (optional), Existing validated instruments for the construct (optional), Expected distributions or prior data (optional), Markets, languages, literacy and accessibility profile (optional)

Typical outputs

What you get back.

Scale specification table covering form, points, labels, direction, midpoint and statistic, Design rationale notes for every contested decision, Analysis compatibility check per construct, Composite measures assessment for any net, index or gap score, Comparability register naming what is locked, what is breaking and the bridging plan, Consolidated review points

Method coverage

What the skill works through.

  1. Why scale decisions get made by habit and discovered at analysis
  2. Defining the construct, and whether one item can carry it
  3. Adopting a validated measure versus writing your own
  4. Unipolar or bipolar: absence or opposite
  5. How many points, and what the evidence actually supports
  6. Labelled, numbered, endpoint-labelled, and why label intensity matters more
  7. Deciding the midpoint, and separating neutral from no opinion
  8. Agreement scales and the acquiescence they invite
  9. Frequency scales and vague quantifiers
  10. Behavioural versus attitudinal measurement
  11. Matching the scale to the analysis planned
  12. Net scores, index scores and difference scores
  13. Changing a scale mid-tracker, and how to bridge it

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

How many points should a rating scale have?

It depends on what the number will be used for, and the evidence supports a range rather than a single answer. Reliability rises with the number of points and plateaus around five to seven on most constructs. Discrimination and sensitivity to small change keep improving modestly beyond that, which matters for trackers and for correlational work. Burden and error rise once the scale exceeds what someone can hold in mind, which bites hardest when it is read aloud or shown on a small screen. So: five to seven for a distribution or a top-box report, seven to eleven where the measure feeds correlation or regression, four or five with short labels when read aloud, and fewer points presented vertically on small screens. Longer formats are often chosen for benchmark comparability rather than for measurement reasons, which is legitimate as long as it is the stated reason.

Should a scale have a midpoint?

Only where a genuine neutral position exists and some respondents hold it, which is true of satisfaction, agreement and comparison against a benchmark, and false of unipolar constructs where there is no middle between none and a lot. Removing a midpoint where one really exists does not manufacture opinion; it pushes neutral respondents to whichever adjacent point is more socially comfortable, usually the positive one. Including one where none exists gives satisficing respondents a resting place. Separately, keep the midpoint distinct from "don't know" or "not applicable": a neutral position is a view and an absence of view is not, and collapsing them makes both uninterpretable.

What is the difference between a unipolar and a bipolar scale?

A unipolar scale runs from none to a maximum of one thing (not at all useful to extremely useful; never to always). A bipolar scale runs from one pole through a neutral centre to its opposite (very dissatisfied to very satisfied). The test is whether the construct has a genuine opposite or merely an absence. Usefulness has an absence: something can fail to be useful without being harmful, so a bipolar usefulness scale invites respondents to endorse a position that does not exist. Getting this wrong has a detectable signature: respondents holding the "none" position cluster on the midpoint or the first negative point, producing a shoulder that looks like mild negativity and is actually absence.

Why are agree/disagree scales a problem?

Because they measure agreement, and agreement carries an acquiescence component: a general tendency to agree with statements, stronger in some populations than others, which inflates every item in the same direction and correlates the whole battery so that it looks like a construct. Construct-specific frames ask about the thing directly with a scale built for it: "how easy or difficult was it to..." rather than "the process was easy: agree or disagree". They outperform agreement frames on interpretability, comparability and acquiescence resistance at no cost in length. Where an agreement battery is unavoidable, reverse some items, keep the battery short, and avoid several same-direction batteries in sequence.

Should I use an existing validated scale or write my own?

Use an existing one when the construct has been studied, when comparability to a literature matters, or when the finding will be scrutinised, because you inherit reliability and validity evidence rather than having to generate it. Write your own when the construct is specific to the client's context, when no validated measure exists, or when the validated instrument is too long and truncating it would break the validation anyway. Three disciplines when adopting one: use it in full and unmodified where possible, disclose any modification and stop claiming the original's properties, and check that the validation population resembles yours.

What is wrong with net scores?

A net (one category minus another) discards the middle of the distribution and the intensity within categories, so two very different distributions produce the same net and a movement in it cannot be decomposed without going back to the underlying data. Nets are also more volatile at small bases than their components, because two proportions each carry sampling error and the net compounds both, which is why nets in small subgroups move for no reason. The rule is not to abolish them but to report the distribution alongside, never let the composite be the only number that travels, state the base below which it is not reported, and watch for movements driven by one component while the headline stays flat.

Are importance-minus-performance gap scores reliable?

Less than either component. A difference score inherits the measurement error of both measures and typically has lower reliability than either alone, and it is usually dominated by whichever component varies more, which is often not the one the client cares about. There is a further problem specific to importance ratings, which is that stated importance discriminates poorly because respondents rate nearly everything important. Report both components alongside any gap, and where the objective is a genuine trade-off, use a designed exercise rather than a rating-based gap.

How do I ask about frequency without vague words?

Avoid "often", "regularly" and "sometimes", which are interpreted differently by different respondents and shift with the base rate of the behaviour, so the same word describes different frequencies for a daily act and an annual one. Prefer absolute frequencies with defined units and windows: "in a typical week, on how many days did you...". Two cautions come with that. Absolute counts inherit all the recall problems of routine low-salience behaviour, so shorten the window until the respondent can reconstruct it, or ask about the most recent occasion instead. And the ranges themselves anchor, because respondents infer the normal range from the options offered.

Can I change a scale in the middle of a tracker?

Only deliberately, with a bridge, and usually the honest answer is not to. The movement introduced by a scale change is indistinguishable from real movement, and the gain from a better measure is normally smaller than the loss of a comparable trend. Where a change is genuinely necessary, bridge it with a parallel run: field both versions in the same wave to random halves of the sample, so the relationship between them is estimated on one sample at one point in time with the halves comparable in every other respect. Report both series for at least the bridging wave. Do not retro-convert history with a factor estimated across two different waves, because that confounds the scale change with real change.

Should I measure behaviour or attitude?

Behaviour, wherever both are available, because it is checkable, more stable, and more predictive of behaviour than attitude is. "Have you recommended this to anyone in the past six months, and to whom" is a different and better measure than "how likely are you to recommend". Attitudes earn their place where behaviour is unobservable, where the attitude precedes a behaviour the study wants to anticipate, where the diagnostic value is in the reasoning, or where the behaviour is too rare to measure in the sample. Whichever is used, the report must say which, because attitudinal measures are routinely presented as though they described behaviour.

How does the scale affect the analysis I can do?

Directly, and it is easier to fix before fielding than after. Means assume interval spacing that ordinal scales do not strictly meet, an assumption widely accepted in practice that should be stated rather than hidden, and it fails harder on short scales and unevenly spaced labels. Top-box percentages need enough points for the box to be a meaningful subset and discard the rest of the distribution. Correlation, regression and factor analysis want more points and suffer badly from ceiling effects, so a measure everyone answers at the top produces a null result that is an artefact. Significance testing on means requires an identical scale across the groups compared. Walk the chain backwards from the chart before finalising anything.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out