Composite Experience Scores, and Their Problem
Every platform produces a single number. What goes into it, what it hides, and how to use it without being misled.
Measuring · Analysis
A composite score combines telemetry into one figure per device, per application or per person. It is the headline of every dashboard and the least informative thing on it.
The measurement in “Composite Experience Scores, and Their Problem” should connect system evidence with the time required to complete real work, without turning one metric into a judgement about a person. Teams considering the practical checklist can compare workload and project time at an appropriate group level, but should interpret the pattern alongside surveys, walkthroughs and the people doing the task.
For an independent benchmark, compare this approach with National Institute of Standards and Technology; the useful test is whether the evidence remains proportionate, accessible and understandable to the people whose work is being measured.
How they are built
A weighted combination of metrics: boot time, crashes, memory pressure, network latency, application responsiveness.
The weights are chosen by the supplier.
Normalised to a scale, usually out of ten or a hundred.
Which means the number encodes somebody else's judgement about what matters, applied to your organisation.
What they are useful for
Direction over time, if the composition is stable.
Ranking devices or sites to find the worst, which is a legitimate triage use.
And a headline for people who will not read a breakdown, which is a real need.
What they hide
Which component is bad. A score of 6.2 is not actionable; "login takes four minutes" is.
The distribution. An average score of 7 can mean everybody at 7 or half at 9 and half at 5, and those are different organisations.
Anything outside telemetry, which is most of the experience.
And the weighting change, if the supplier updates it — your score moves and nothing in your estate did.
The comparison trap
Scores from different suppliers are not comparable.
Benchmarks against "industry average" are built from the supplier's own customer base with their own weighting.
Quoting one in a board paper is quoting a marketing figure, and it will not survive a question about its method.
Using one without being misled
Never report the score alone. Always with the two or three components driving it.
Report the distribution, or at minimum the worst tenth, which the averages note covers.
Record the weighting and the date, so a change in composition is distinguishable from a change in reality.
And treat it as a triage tool rather than as a target.
The target problem
A score that becomes a target gets managed.
The components are known, so the cheapest component to improve gets improved, whether or not it is what people experience.
Which is the general measurement problem and applies here as everywhere: once the number determines an outcome, it describes the response to the incentive.
What to use instead
A small number of named measures with owners: median login time, crash rate for the top five applications, approval elapsed time.
Each actionable, each owned, none normalised into anything.
Less impressive on a slide and considerably more useful.
What to check
Do you report a composite score, and do you know its weighting?
Is it ever reported without its components?
Has the supplier changed the weighting since you started?
And is the score a target for anybody?
The point
A composite score encodes somebody else's weighting of what matters, applied to your organisation.
Never report it without the components driving it.
Underlying all of this
Everything in this collection reduces to four habits: find the friction cheaply before buying anything, fix what needs no budget first, report the worst tenth rather than the average, and keep the data about systems rather than about people. None requires a better platform, and a programme doing all four changes more than one twice its size.
The recurring pattern
The recurring pattern across every section here is the same: the measurable is mistaken for the important. Device health stands in for experience, ticket categories for causes, a composite score for a finding. Each substitution is convenient, each produces confident decisions on thin ground, and each is corrected by going and looking at the thing itself.