07.05Qualitative AnalysisAvailable

Sentiment and Emotion Analysis

Score sentiment under validity controls: agreement measured against human labels, a cause attached to every number, and refusal where the text cannot carry it.

Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.

What this skill does

The method, encoded.

Standalone sentiment scoring is one of the weakest things routinely sold as analysis. A score compresses text into a number whose meaning nobody defined, produced by a classifier nobody validated, reported on a base nobody checked, and narrated as movement when it moves inside its own error. It answers a question the business did not ask, in place of the one it did.

This skill says that plainly, and then does the job properly, because sentiment with controls is genuinely useful: it prioritises what to read, segments a large corpus, and makes a coded dataset navigable. It is never a finding on its own.

The controls are specific. Screen the corpus for the conditions that defeat classification: sarcasm, negation, comparison, conditionals, mixed sentiment in one response, code-switching, translation, domain vocabulary that inverts ordinary polarity, and text too short to carry meaning. Build a human-labelled gold set and report agreement per class and per segment alongside the human baseline, because an overall accuracy figure hides a classifier that is near random on the short items making up a third of the corpus. Pair every published figure with the content codes that explain it. On a tracker, publish the noise band first, and report movement inside it as no change without offering an explanation.

Best used for

  • Attaching valence to an already-coded frame so it reads as problems and strengths
  • Prioritising which of thousands of comments a human should read first
  • Aspect-level sentiment on distinct parts of a service
  • Validating or challenging an existing automated sentiment feed
  • Detecting genuine step changes in a text stream with a noise band set in advance
  • Segmenting a large corpus before qualitative analysis of a sampled subset

Typical inputs

What you give it.

Complete untruncated text with an identifier on every item, Content codes from prior coding, The question, channel or elicitation that produced the text, A human-labelled reference sample, or the ability to produce one, Closed satisfaction or recommendation measure on the same respondents (optional), Aspect or entity annotations (optional), Prior period data using the same classifier and frame (optional), Domain vocabulary or category glossary (optional)

Typical outputs

What you get back.

Failure-condition screening profile with incidence per condition, Human-labelled gold set with human baseline agreement, Validation table with agreement per class, per segment and per text-length band, Sentiment by content code with drivers and verified verbatims, Aspect-level sentiment preserving mixed valence, Reportability and suppression log naming the gate each suppressed cell failed, Confidence stated per segment rather than overall, Tracker noise band with a movement classification rule, Emotion classification with its own agreement figures and reliability caveat

Method coverage

What the skill works through.

  1. What a sentiment score is, and what it is not
  2. Why sentiment sits downstream of coding, never instead of it
  3. The conditions that break sentiment classification
  4. Sarcasm, negation, comparison and conditionals
  5. Mixed sentiment in a single response
  6. Domain vocabulary where a negative word is neutral
  7. Short text, and why it is where sentiment fails
  8. Building a human-labelled reference set and measuring agreement
  9. Reporting confidence per segment rather than overall
  10. Pairing every score with its cause
  11. Emotion classification and its weaker reliability
  12. Sentiment trackers: establishing a noise band before narrating movement
  13. When to refuse to report a sentiment figure at all

Download

Free skill. One file.

Enter your email once. Every skill you download after that takes a single click.

How to install

Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.

Download skill

Questions

Common questions.

Is sentiment analysis accurate?

It depends entirely on the corpus, and the only honest answer comes from measuring it on yours. A classifier's published accuracy was established on someone else's text, in another domain, at another length. Build a human-labelled sample of 200 to 400 items, measure agreement per class and per segment, and report the human baseline alongside: if two trained labellers agree on 78 percent of items, a machine reporting 91 percent is not outperforming them, something is wrong with the comparison.

Why does sentiment analysis get sarcasm wrong?

Because surface polarity inverts and the cue is contextual rather than lexical. It is not the only failure of this kind: negation scope, comparison ("better than the old one" says nothing absolute), conditionals ("if they fixed it I would be delighted" reads positive and is a complaint), and reported speech all defeat classification the same way. Screen a sample for each condition, record the incidence, and disclose the direction of bias with your results.

What do I do when one response contains both praise and complaint?

Preserve both. "The staff were lovely but I waited fifty minutes" is not a neutral experience, and averaging it into neutral deletes two findings at once. Score at aspect or clause level so opposing valences attach to different things, count mixed items as their own category, and report the mixed proportion, because a rising one usually means your frame is losing resolution.

Why can I not just report one sentiment number for the business?

Because it aggregates incommensurable things. A billing complaint and a compliment about a shop assistant do not average into a meaningful quantity, and the figure moves whenever the channel mix moves, independent of how anyone feels. Report by content code, by aspect and by segment, each with its own base and confidence.

My sentiment score moved three points. Is that a finding?

Usually not. Compute a noise band from the classifier's measured error plus sampling variation on your period base, publish it before the next wave lands, and report any movement inside it as no change, in those words, with no explanation attached. Narrating small movements teaches an organisation that the metric is meaningful at a resolution it does not have. If the band is wider than the movements the business cares about, the metric cannot serve as a tracker.

Can I use sentiment analysis instead of coding open ends?

No. Sentiment is downstream of coding, not a replacement for it. Valence with no subject attached cannot be acted on, because nothing follows from it: "62 percent negative" tells nobody what to change. "Negative sentiment on delivery, driven by the delivery-window code" does. Any sentiment output with no content codes behind it should not be produced.

How reliable is emotion classification?

Materially less reliable than valence. The categories are less distinct, human agreement is lower, expression varies by culture and language, and the label set is a theoretical choice rather than a natural fact. Use a small set defined for your corpus, build a separate gold set, measure agreement separately, and suppress any category that falls below your valence threshold. Absence of an emotion word is not evidence the emotion is absent, and its presence is a statement, not a diagnosis.

Should I run sentiment on translated text?

Only for triage, and marked as translation-dependent throughout. Translation systematically normalises tone: sarcasm, understatement, politeness formulas and intensifier norms are the first things lost. Sentiment on translated text measures the translator as much as the participant, and cross-market comparison of translated sentiment produces differences of ten points or more that have nothing to do with satisfaction.

What base do I need to report a sentiment percentage?

At least 100 items in the specific reporting cell, with 30 as an absolute floor below which no percentage is permitted at all. The threshold applies to every subgroup and every period separately, not to the corpus total. Cells that fail are reported as counts and verbatims, and shown as suppressed with the reason, because an empty dashboard cell reads as zero.

Does neutral sentiment mean customers are satisfied?

No, and this is the most common misreading. Neutral means no clear valence was expressed, which includes factual description, procedural answers and items that say nothing evaluative at all. The complement of negative is not satisfaction, it is everything else. Keep "no sentiment expressed" as a separate category rather than folding it into neutral, or the middle of your distribution becomes uninterpretable.

Research where people already are.
Analyse it where you already work.

Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.

New Report on SA Gambling Impact
Check It Out