Skip to content
Sections
All notes

All notes · Reading

What a Score Cannot Tell You

The limits of quantified experience, stated plainly, so that the programme is not asked to answer questions it cannot.

Reading · Analysis

Experience measurement narrows uncertainty about specific, countable things. Several adjacent questions get asked of it and are outside what it can support.

The measurement in “What a Score Cannot Tell You” should connect system evidence with the time required to complete real work, without turning one metric into a judgement about a person. Teams considering task time tracking can compare workload and project time at an appropriate group level, but should interpret the pattern alongside surveys, walkthroughs and the people doing the task.

For an independent benchmark, compare this approach with National Institute of Standards and Technology; the useful test is whether the evidence remains proportionate, accessible and understandable to the people whose work is being measured.

Whether people are productive

Fast devices do not produce output and slow ones do not prevent it.

The relationship exists and is weak, mediated by everything else about the job.

Any claim of the form "we improved scores therefore productivity rose" is unsupported, and the figures circulating for this are vendor estimates with assumptions buried in them.

Whether people are happy

Experience is one input to how work feels, alongside the manager, the workload, the pay and the colleagues.

A programme that fixes logins will not move engagement, and promising that it will is how a programme gets judged against a target it cannot reach.

Why somebody is having a bad time

The score says the device is poor. It does not say whether that is the hardware, the configuration, the applications, the network or what the person is trying to do.

That requires looking, which is the segmenting note's argument.

Whether the fix was worth it

Cost is knowable. Benefit is time saved multiplied by an assumed value of time, and the assumption carries the result.

Which is why the value note argues for stating time saved rather than converting it to money, unless your finance function provides the rate.

What people would want instead

Measurement describes what exists. It cannot tell you what people would prefer, which requires asking.

A programme optimising only measured dimensions improves those dimensions and may miss what matters most, exactly as occupancy sensing removes the quiet corner nobody counted.

Anything about an individual

Covered in its own note and worth repeating here: a per-person score describes the person's equipment and circumstances, not the person.

It will be read as describing the person anyway, which is why it should not exist.

What it does support

That this cohort waits longer than that one.

That this application fails more than it did last month.

That this change reduced this measure by this much for this group.

Narrow, checkable claims. The programme's credibility rests on making those and refusing the others.

The sentence worth having ready

"We can tell you how long things take and how often they fail. We cannot tell you whether people are productive, and nobody can."

Said early, it sets expectations. Said after a board paper overclaims, it is damage control.

What to check

Has your programme been asked to show a productivity effect?

What did you say?

Are the claims in your reporting checkable?

And does anybody expect this work to move an engagement score?

The point

The programme can say how long things take and how often they fail.

It cannot say whether people are productive, and nobody can.

Underlying all of this

Everything in this collection reduces to four habits: find the friction cheaply before buying anything, fix what needs no budget first, report the worst tenth rather than the average, and keep the data about systems rather than about people. None requires a better platform, and a programme doing all four changes more than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the measurable is mistaken for the important. Device health stands in for experience, ticket categories for causes, a composite score for a finding. Each substitution is convenient, each produces confident decisions on thin ground, and each is corrected by going and looking at the thing itself.