Skip to content
Sections
All notes

All notes · Fixing

Measuring Whether the Fix Worked

Most improvements are claimed rather than demonstrated. Four steps that make the claim survive a question.

Fixing · Procedure

A change was made, the dashboard looks better, and the improvement is announced. Whether the change caused it is a separate question that usually goes unasked.

The recommendations in “Measuring Whether the Fix Worked” need visible ownership, review time and a way to show whether the change reduced effort for the affected group. An organisation can use Monitask to coordinate that implementation work and compare workloads, without treating hours or activity as a complete measure of digital employee experience.

For an independent benchmark, compare this approach with Nielsen Norman Group; the useful test is whether the evidence remains proportionate, accessible and understandable to the people whose work is being measured.

Before

Measure the specific thing, for the specific population, for long enough to know its ordinary range.

Two to four weeks, unless the measure is stable.

Write the number down, with the date and the definition.

This step is the one skipped, and without it everything afterwards is an assertion.

During

Change one thing.

If three changes land in the same week, you will never know which worked, and the next similar decision has no evidence behind it.

Where several changes are necessary, stagger them by a fortnight.

After

Measure the same thing, the same way, for the same population.

Long enough to clear the novelty and the ordinary variation.

And compare against the ordinary range rather than against the single prior value — a move within normal variation is not an effect.

The control

Where you can, roll out to half the population first.

The unchanged half is a control, and it removes the seasonal and organisational effects that otherwise contaminate everything.

This is cheap, rarely done, and the difference between a claim and a finding.

Where a control is impossible, say so when reporting.

What contaminates the comparison

Hardware refresh running in parallel.

A quarter-end or holiday period either side.

Agent deployment widening, which the baselines note covers.

Anything else changed by anybody, which in a large organisation is a lot.

List what else happened in the window and say whether it could explain the result.

Reporting it honestly

State the population, the before and after, the window, and what else changed.

"Median login for the 1,400 users on the old image fell from 3m10 to 1m25 over six weeks; no hardware changed in that group."

That survives a question. "Experience improved 23%" does not.

When it did not work

Say so.

A programme that reports only successes is not believed about any of them.

And a failed fix is information: it rules out a cause, which is worth having.

What to check

Do you have a before figure for your last change?

Did you change one thing or several?

Was there a control group?

And has your programme ever reported something that did not work?

The point

Measure before, change one thing, measure after against the ordinary range, and say what else happened in the window..

Underlying all of this

Everything in this collection reduces to four habits: find the friction cheaply before buying anything, fix what needs no budget first, report the worst tenth rather than the average, and keep the data about systems rather than about people. None requires a better platform, and a programme doing all four changes more than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the measurable is mistaken for the important. Device health stands in for experience, ticket categories for causes, a composite score for a finding. Each substitution is convenient, each produces confident decisions on thin ground, and each is corrected by going and looking at the thing itself.