Slower is not broken. Broken raises a ticket. Slower just sits there.
We put an AI-based reporting tool into a group of radiology practices. Live worklists, real cases, a hard go-live. The program reported an efficiency gain across the reader population, and that gain was real. This is not a story about a deployment that failed.
Inside the gain sat a group of readers whose measured efficiency had gone negative. Not flat. Negative. They were getting through less work per hour than before we arrived, and they had been that way for some time.
Not one of them had raised it.
We eventually stood up a dedicated workstream for that group and ran it for months. The workstream is not the interesting part. How long they stayed invisible is, and why.
The average did what averages do
A gain in the middle of a distribution plus a tail at the bottom is still a gain. The number was not wrong. The reporting answered the question we had asked it, which was whether the population got faster, and the answer was yes.
That is a fine question for a steering committee and a useless one for an operator. Did the population get faster, and who did this make worse, are two different questions with two different owners. Only one of them was on a dashboard.
I have built the aggregate first every single time. It is what leadership asks for, it is cheap to produce, and it is right often enough to feel sufficient. The per-account risk score I built at Ceridian existed because the aggregate satisfaction number had already failed me. I did not start there. I got there.
Two things have to be true for a cohort to disappear
The first is that the system has no way to represent the failure. There is no error state for slower. No threshold, no alarm, no field, no ticket type. A reader taking longer on a study than they used to has generated no event anywhere in the stack. Every component returned success.
The second is that the person living it has a reason not to say so. This is the half operators skip, and it is the half that does the damage. Being slower on a new tool does not read as a product defect to the person being slow. It reads as a personal deficiency. Nobody opens a case to report that they have gotten worse at their own job in front of their partners.
Both conditions hold at once in every clinical deployment I have run. The silence is structural.
Which means the silence tells you nothing about the size of the tail.
Measure the sign, not the level
The instrument that finds this is not a better dashboard. It is a different one, and you build it on purpose.
A utilization dashboard reports level: how much of the tool is in use, by how many people, against a target. Level is what everyone builds, because level is what a rollout is judged on. Level cannot see our cohort at all. Those readers were using the tool constantly. They were using it and getting slower, which registers on a utilization chart as success.
What finds them is per-reader change in direction. Take each individual's throughput before and after, and alarm on the sign of the difference rather than its value. You are not asking how much better someone got. You are asking whether the arrow points down, for this person, at all. Count the arrows pointing down. That count is your workstream.
Most operators tell me they cannot do this because their before-data is thin. Usually it is. There are two ways around that. Use a rolling internal baseline: each person measured against their own first weeks on the tool, once the initial learning spike has settled out. Or run a cohort comparison against the group that has not onboarded yet, which any staged rollout hands you for free.
And if you have no before at all, measure slope instead of delta. You do not need a starting point to know which way somebody is travelling. A curve falling over six weeks is the same alarm whether or not you know where it started.
Small technical change, large organizational one, because it produces a number nobody wants: how many people your program made worse. Somebody has to be willing to publish it.
What the tail was actually telling us
Here is the part I did not expect.
At one practice, an exam capture efficiency curve went negative and then came back. It recovered when we corrected an underlying defect. Nothing about that recovery involved training, coaching, or persuading anyone. We fixed a thing that was broken and the curve turned.
The people at the bottom of the distribution had not failed to adapt. They were standing closest to the problem.
Which reframes the cohort. The negative tail was not a population that needed remediation. It was the earliest clean read on a defect we had not found yet.
That changes what the workstream is for. You are not staffing a rescue. You are staffing an instrument, and the people in it are your best sensor.
The same move, in a business with no radiologists in it
At Ceridian I built a Customer Risk Score for exactly this reason. We had an aggregate satisfaction number and it was hiding accounts. So we stopped reading the aggregate, scored every account individually, and staffed the bottom of that list as its own book of work. NPS moved from -52 to +20 across 4,000+ customers on a $400M P&L.
Different industry, different decade, same three moves: stop reading the mean, score the individual, staff the bottom.
The clinical version carries one more piece of evidence. In our own population, radiologists using drafting recommend the platform at 67 percent against 46 percent for reporting-only users. That is not an average telling you the deployment is going well. It is a distribution telling you it is going differently for different people, which is the only useful thing a number about adoption can tell you.
Where this breaks
Two places, and I have hit both.
On a small enough population, per-person measurement stops being instrumentation and starts being surveillance. There is a real threshold below which publishing individual directional data costs more trust than the signal is worth, and it is lower than most operators think. If you cannot name the cohort without naming the person, you do not have a cohort.
The design move that buys most of the safety back is to separate measurement from reporting. Measure at the individual level, because that is where the signal is. Report at the cohort level, because a cohort is what leadership can actually act on. And show a person their own number before anyone above them sees it. A reader who finds out from a dashboard that leadership has been tracking their decline has told you everything they are ever going to tell you about their workflow.
And some negative curves are a learning slope that resolves on its own. Staffing every one burns capacity you needed elsewhere. The test I use is whether the curve is still moving. A curve climbing back, however slowly, is a person adapting. A curve gone flat at the bottom is a defect wearing a person's name.
Go look at your distribution. The tail is not going to raise its hand.
Also in the series · The Two Scorecards
First published on LinkedIn, August 2026.