05.01Quantitative AnalysisAvailable

Descriptive Analysis

Frequencies, percentages, means and distributions done properly: every figure on a stated base, every summary chosen from the distribution, nothing overstated.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Descriptive analysis is the first thing done to a dataset and the most underrated skill in quantitative research. The arithmetic is easy, which is why the errors are rarely arithmetic. They are denominator errors: a percentage calculated on the people who were asked a filtered question, then reported as though it described everybody. They are summary errors: a mean quoted for a distribution so skewed that almost no respondent sits near it. They are omission errors: figures with no base sizes, several resting on fewer than fifty people. And they are volume errors, because a descriptive run produces hundreds of numbers and reporting all of them is not analysis.

This skill encodes the professional method. It builds a base register before it calculates anything, so every denominator can be reconciled to the routing. It reads each distribution before choosing how to summarise it, and says when the median, the mode or the whole distribution is the honest answer. It handles multi-response, top-box and net scores as declared conventions rather than defaults. It applies small-base thresholds without negotiation, discloses weighting and effective bases, and refuses false precision.

It describes differences. It does not test them, and it says so.

Best used for

  • Producing toplines and data tables from a fielded survey
  • Establishing the headline figures a report will be built on
  • Getting denominators right on questionnaires with filters and routing
  • Deciding whether a mean, a median or the full distribution is the honest summary
  • Auditing percentages somebody else produced, particularly their bases
  • Preparing verified, correctly based inputs for cross-tabulation and testing
  • Reporting small subgroups honestly rather than dropping them

Typical inputs

What you give it.

Prepared respondent-level dataset, one row per respondent, Questionnaire as fielded, with question wording and scale labels, Routing and filter logic, Data dictionary or codebook with value labels and missing-value codes, Research objectives, Weighting variable and weighting specification (optional), Sample and fieldwork documentation (optional), Cleaning log (optional), Previous wave tables and conventions (optional)

Typical outputs

What you get back.

Conventions and disclosure block covering bases, weighting, rounding and scale conventions, Base register reconciling every denominator to the routing, Frequency tables with base description in words and base size on every figure, Summary statistic table with distribution shape and the reason each statistic was chosen, Multi-response tables labelled, with mean selections per respondent, Top-box and net scores with their convention stated, Small-base register listing every suppressed subgroup and why, Findings summary ordered by research objective, List of figures produced but not reported, with the selection rule, Statement of what could not be established

Method coverage

What the skill works through.

  1. What descriptive analysis is, and where it stops
  2. Base discipline: every percentage has a denominator
  3. Filtered questions, "of those who", and rebasing to the total sample
  4. Multi-response: why percentages exceed 100 and how to label it
  5. Reading a distribution before summarising it
  6. When the mean is the wrong statistic
  7. Means on ordinal scales: a convention, not a measurement
  8. Top-box and net scores: usefulness, information loss, stating the convention
  9. Small bases and the reporting thresholds
  10. Don't know, refused and not asked: in the base or out
  11. Weighted figures and the effective base
  12. Rounding, false precision and totals that do not sum to 100
  13. Selecting figures against research objectives instead of dumping them
  14. Describing a difference is not claiming one

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

What is a base in survey analysis, and why does it matter so much?

The base is the group of people a percentage is calculated on, described in words and counted. It matters because the same number means entirely different things depending on it: 58% of past-month users is not 58% of customers, and it is not 58% of the population. Every reported figure needs both the base description and the base size next to it, not in a footnote.

Why do my survey percentages add up to more than 100?

Almost always because the question was multi-response and people could pick more than one answer. Percentages are calculated on respondents, not on selections, so the column sums above 100. Say so in the table title and report the mean number of selections per respondent, because a list averaging four selections behaves nothing like one averaging 1.3.

Why don't my percentages add up to exactly 100?

Diagnose before you adjust. A total of 99 or 101 on a handful of categories is rounding, and gets a standard note. A total of 96 or 104 is not rounding: it usually means a base error, a missing category, an overlapping code frame, or a multi-response question being treated as single. Never nudge a category to make a column balance, because that destroys your best early warning.

Should I report the mean or the median?

Look at the distribution first, then decide. Where the data is skewed, which is normal for income, spend, transaction counts, tenure and visit frequency, the mean sits above almost everyone and the median describes the typical respondent. Report the mean as well when the total is what matters, such as revenue or volume, because the mean and the total carry the same information. Where the distribution has two peaks, no single average is honest and the distribution itself is the finding.

Can I calculate an average on a five-point rating scale?

You can, and it is a convention rather than a measurement. A mean assumes the gap between "strongly disagree" and "disagree" equals the gap between "neutral" and "agree", which nobody has tested. Mean scores are genuinely useful for tracking small movements over time, and should always be shown with the distribution beside them and the convention stated.

What is a top-box or top-two-box score, and what does it hide?

It is the proportion choosing the most positive one or two points on a scale. It communicates well and it discards the shape of the distribution: two quite different sets of answers can produce the same score, and a net score that is stable across waves can conceal both ends growing at once. State the convention in the same place as the number, never compare a top-two-box figure with a top-three-box figure, and keep the full distribution in the deliverable.

What is the minimum base size for reporting a percentage?

Below 100, show the base, flag it, and read it directionally only. Below 30, do not report percentages at all: report counts. A useful check is that one respondent is worth 100 divided by the base in percentage points, so on a base of 40 a single person moves the figure 2.5 points. When a subgroup is too small, say which subgroup, give its base, and say what would be needed to report it, rather than quietly dropping it.

Should "don't know" be included in the base?

Decide once, apply it everywhere, and disclose it. For opinion, awareness and knowledge questions a don't know is usually a substantive answer and belongs in the base. Excluding it silently converts "of those who had a view" into "of everyone", which is a base error wearing different clothes, and it commonly moves a headline by four points or more. "Not asked" is different: those people were routed away and never enter the base at all.

How do I report weighted survey data honestly?

Label the figures as weighted, name the weighting scheme, show the unweighted count, and show the effective base. A weighted count is not a number of interviews and must never be presented as one. Where weights vary a lot the effective base can be much smaller than the achieved sample, and it is the effective base that governs small-base thresholds and everything downstream.

Should I report percentages to one decimal place?

Rarely. A decimal claims a precision that sampling does not provide, and on any base under 1,000 a tenth of a point is finer than a single respondent, so it is meaningless by construction. Whole percentages are the default, and precision should never do rhetorical work that the method cannot support.

How do I know which numbers to report out of the hundreds a descriptive run produces?

Write the output plan before running anything: list the research objectives and, against each, the specific figures that would answer it. Report those, plus anything genuinely surprising, plus the profile figures a reader needs in order to know who was surveyed. Everything else goes to an appendix, with a note saying what was produced and not reported so the selection hides nothing. The rule has to be written before the figures are seen, or selection becomes cherry-picking.

Can descriptive analysis tell me whether two groups are different?

No. It can show that one group answered 44% and another 37%, with both bases visible, and that is a description. Saying the groups genuinely differ is a claim about reliability and needs a statistical test behind it. Describing a difference is not claiming one, and a good descriptive output makes sure a reader cannot mistake the first for the second.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out