18 January 2026
Accuracy is a flattering metric for rare cancels
A March sitting brought a model with 97% accuracy and a save queue that was always empty. Almost nobody cancelled in a given week. Predicting “stays” was a way to look busy. The Feature Hygiene clinic found nothing illegal in the columns; the scoreboard itself was the vanity.
We keep three numbers on the wall that are less flattering. First, precision among the accounts a human will actually phone — if that list is 200 people, how many were truly on a path to cancel. Second, lead time: days between a honest mark and the cancel event, for the ones we did mark. Third, contamination: how many marks were themselves caused by last quarter’s outreach.
Calibration is not a vibe
If a score of 0.4 is supposed to mean four in ten, we count. When it means one in ten, we say so in the detection brief and we stop using the score as a queue rank. Ranking can still be useful; pretending it is a probability is how finance gets a second tool.
Accuracy can remain in an appendix if someone insists. It does not get a slide of its own in our reviews.