A score between nought and a hundred is worthless if nobody can say why it is that number. We rebuilt scoring around explanation and lost a little accuracy doing it.

Nobody trusted the old score

The original model was more accurate than what replaced it. It was also almost entirely unused, because a salesperson given a 73 with no explanation has no basis for acting differently than they would have anyway.

An unused accurate score is worth less than a used approximate one.

Every point is attributable

The new model is additive and every contribution traces back to a signal you can see. Opening the score shows the signals, their weights and their direction.

  • Each signal is named in plain language, not as a feature id
  • Signals can be switched off per workspace when they do not apply
  • A score that moves reports which signal moved it and when
  • Signals with thin data are shown as thin rather than silently smoothed

The accuracy trade

Constraining the model to additive, explainable signals cost a few points of precision against our holdout set. We took it deliberately.

The interesting result is that end-to-end outcomes improved anyway, because the score started being acted on. Model accuracy was never the binding constraint.

What we would do differently

Build the explanation first and fit the model to it, rather than building a model and trying to explain it afterwards. Retrofitting explanation onto an existing model is much harder than it sounds, and the explanations you get are worse.

We use cookies

We use strictly necessary cookies to run the site, and optional cookies to understand usage and measure marketing. You can change your choice at any time. Cookie policy