- Home
- Research skills
- Driver Analysis
Driver Analysis
Find which attributes are most strongly associated with an outcome, test whether the ranking is stable, and report associations rather than drivers.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
Every organisation running a satisfaction or loyalty study eventually asks which of the twenty things it measures actually matter. Answering it badly is easy. Ask people directly and everything comes back important. Correlate each attribute with the outcome one at a time and the ranking reflects how correlated the attributes are with each other. Run a regression on twenty attributes that all correlate at 0.6 and the coefficients become unstable to the point of arbitrariness: drop one and the order changes, add respondents and a sign flips. Then the list goes into a deck headed "drivers", and money is spent on attribute three because it is third.
This skill handles the naming problem first. The technique is called driver analysis and it does not identify drivers; it identifies statistical associations, and a cross-sectional design does not license the causal reading. It then covers the choice between correlation-based, regression-based, relative-importance and structural approaches; multicollinearity and how to test whether a ranking survives perturbation; whether the outcome has enough variance to model; why an attribute absent from the questionnaire cannot appear; the importance and performance quadrant and its misuses; and a reporting standard that keeps the causal claim out.
Best used for
- Ranking service or product attributes by association with satisfaction or loyalty
- Prioritising improvement effort when everything cannot be improved
- Comparing what respondents say matters against what covaries with the outcome
- Building an importance and performance prioritisation view
- Auditing a supplied driver model for stability and overreach
- Bringing a causal claim back within what the design supports
Typical inputs
What you give it.
Defined outcome variable with wording, scale and direction, Attribute battery as fielded, with wording and scale direction per item, Respondent-level dataset with outcome and attributes on the same base, Modelling base after missing-data treatment, The decision the analysis informs, Stated importance ratings for the same attributes (optional), Performance ratings on a comparable scale (optional), Linked behavioural or transactional outcome data (optional), Previous waves of the same model (optional), A prior structural model of the relationships (optional)
Typical outputs
What you get back.
Model specification block with outcome, attribute set, base and approach, Attribute inventory with direction, missingness, correlation and collinearity measures, Derived importance table with stability results per attribute, Stated versus derived importance comparison with table-stakes attributes annotated, Importance and performance view with reference points and boundary uncertainty, Model fit note with base, attribute set and reading of the unexplained share, The association statement placed alongside the ranking, Statement of attributes not measured and the omitted-variable risk
Method coverage
What the skill works through.
- The naming problem: driver analysis does not identify drivers
- Choosing the outcome, and why different outcomes give different rankings
- Does the outcome have enough variance to model?
- Auditing the attribute battery before modelling it
- Stated importance against derived importance, and why they disagree
- Table stakes: why the best-performing attribute looks least important
- Correlation-based, regression-based, relative-importance and structural approaches
- Multicollinearity, unstable coefficients and arbitrary rankings
- Testing whether a ranking survives resampling
- Reading model fit honestly
- The importance and performance quadrant, and its three misuses
- What a driver absent from the questionnaire does to the model
- The reporting standard and the causal firewall
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
What is key driver analysis?
A family of methods that relates a set of measured attributes to an outcome such as satisfaction, loyalty or intention, and ranks the attributes by how strongly they are associated with it. The important qualification is in the name: it identifies statistical associations, not causes. Attributes and outcome are usually measured at the same moment from the same respondent, so temporal order is unknown, reverse causation is unexcluded and confounding is untreated.
Why do stated importance and derived importance disagree?
Because they measure different things, and the disagreement is informative rather than a problem to fix. Stated importance reflects what people believe matters and what is acceptable to say matters; it over-reports price, reliability and safety. Derived importance reflects what covaries with the outcome in this sample, and it favours whatever is caught up in the general favourability running through all the ratings. The most useful cell is low stated and high derived, where something matters that people do not articulate. The most dangerous is high stated and low derived, which usually means the attribute is delivered uniformly well and therefore cannot discriminate.
Why is my best-performing attribute ranked least important?
Because a model detects an attribute through its variation, and an attribute everyone rates highly has almost none. Security, safety and basic reliability routinely rank near the bottom of derived importance for exactly this reason. They are hygiene factors: they do not distinguish satisfied customers from dissatisfied ones while they are working, and they would dominate the model if they failed. Annotate them explicitly, because an instruction to divest from them can be very expensive.
What does multicollinearity do to a driver analysis?
Survey attribute batteries typically correlate at 0.4 to 0.7 with each other. When predictors share that much variance, the regression cannot determine how to split the shared association between them, so individual coefficients become unstable: standard errors inflate, signs can flip, and dropping one attribute reorders the rest. The ranking then depends on which attributes happened to be in the battery. Report a collinearity measure per attribute, and test the ranking by resampling rather than trusting the point estimates.
How do I know whether my driver ranking is stable enough to report?
Perturb it. Re-estimate on bootstrap resamples and record how often each attribute lands in the top three; split the sample in half and compare; drop an attribute and see whether the rest reorder. If the leading attributes hold their positions in the large majority of resamples, the ranking is usable. If an attribute moves between second and ninth depending on the resample, there is no ranking to report, however precise the point estimates look, and the honest output is a grouped set rather than an ordered list.
Can driver analysis tell me what to fix?
It can tell you which of the things you measured are most strongly associated with the outcome, which is a real and useful input. It cannot tell you that fixing one of them will move the outcome, because the design does not establish causation, and it cannot see anything that was not in the questionnaire. It also says nothing about cost, feasibility or how much room for improvement exists. Prioritisation needs all of that, and it needs a person.
What if an important driver was not asked about?
It cannot appear in the model, and its association is absorbed by whichever measured attributes correlate with it. If price was not asked and price is what matters, the model will credit value perception or fairness instead, and the output will look exactly the same as a correct one. This is the largest and least visible risk in the method. List what was not measured alongside the ranking every time, and where a known material attribute is missing, say the analysis cannot answer the question.
How much of the outcome should a good driver model explain?
There is no target, and chasing a high figure usually means adding items that duplicate the outcome. A substantial part of any figure from a single-questionnaire model is the general favourability that colours all of a respondent's ratings, so a high value is not evidence that the business understands its customers. A low value often means the outcome depends on things the study did not measure, which is a genuine finding rather than a defect.
What is the importance and performance matrix used for, and what goes wrong with it?
It plots derived importance against current performance to identify where improvement effort is best placed. Three things go wrong. It invites the reading that raising an attribute's performance will move the outcome in proportion to its importance score, which the design does not support. Attributes near the axis crossing are not distinguishable from each other, so the quadrant boundary gets treated as a decision line when it is not. And well-delivered hygiene factors sit in the low-importance region precisely because they are working, so the matrix can appear to recommend abandoning them.
Can AI run driver analysis reliably?
It applies the diagnostics consistently and does not get bored of running the stability checks, which is a genuine advantage. The risks are specific: producing a ranked list with no model behind it, presenting ranks that the data will not support, letting causal language reappear in a heading or an executive summary after being removed from the analysis, and converting importance shares into projected gains as though they were rates of return. Used properly, AI should be required to report model fit, base and attribute set with every ranking, attach a stability result, state what was not measured, and place the association statement in the same document as the numbers.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
