A score between nought and a hundred is worthless if nobody can say why it is that number. We rebuilt scoring around explanation and lost a little accuracy doing it.
Nobody trusted the old score
The original model was more accurate than what replaced it. It was also almost entirely unused, because a salesperson given a 73 with no explanation has no basis for acting differently than they would have anyway.
An unused accurate score is worth less than a used approximate one.
Every point is attributable
The new model is additive and every contribution traces back to a signal you can see. Opening the score shows the signals, their weights and their direction.
- Each signal is named in plain language, not as a feature id
- Signals can be switched off per workspace when they do not apply
- A score that moves reports which signal moved it and when
- Signals with thin data are shown as thin rather than silently smoothed
The accuracy trade
Constraining the model to additive, explainable signals cost a few points of precision against our holdout set. We took it deliberately.
The interesting result is that end-to-end outcomes improved anyway, because the score started being acted on. Model accuracy was never the binding constraint.
What we would do differently
Build the explanation first and fit the model to it, rather than building a model and trying to explain it afterwards. Retrofitting explanation onto an existing model is much harder than it sounds, and the explanations you get are worse.