Skip to content

All postsEngineering

A Confidence Score Is Not a Probability of Being Right

The number a model reports next to its answer is not the chance that it is correct, and the segments it is most reliable on are where its worst errors hide. We replaced confidence thresholds with a licence per segment, earned on reviewed cases and revoked automatically.

February 11, 20256 min readWritten by David Grayhurst

A financial controller put an invoice on the table. The system had coded it to the wrong cost centre and posted it without review. Next to the decision, the log showed a confidence of 0.97. "If it was 97% sure," she asked, "why was it wrong?"

It was never 97% sure of anything. The number was a score the model produced about its own output. Across many invoices it correlated with being right, roughly the way a person's feeling of certainty does. It was not a measured probability, and for this supplier, after last month's cost-centre reorganisation, it was not even a useful guide.

That conversation ended our use of confidence thresholds for straight-through processing. What replaced them is less elegant and much easier to audit.

Fluent systems sound sure

The problem is well documented. In "On Calibration of Modern Neural Networks" (ICML 2017), Guo and colleagues showed that modern deep networks tend to be overconfident: reported confidence runs ahead of accuracy. Work on language models points the same way. Xiong and colleagues, at ICLR 2024, found that models asked to state their confidence tend to overstate it, clustering at high values whether or not they are right.

In operations the effect is sharper, for a structural reason. The cases a model gets wrong usually look normal: a familiar supplier with a new product line, a clean invoice coded to a cost centre that was reorganised last month. Nothing in the input signals that the world has changed, so nothing in the model's self-assessment changes either. Confident errors are not a rare tail. They are how a fluent system fails when the business changes quietly around it.

Straight-through is a licence per segment

We do not give a system a threshold. We give each segment a licence. A segment is a supplier, or a group of small suppliers, crossed with document type and entity. It is licensed for straight-through processing only when three conditions hold on the client's own reviewed cases: a minimum run of consecutive cases checked by a person, zero errors among them in any field that moves money, and every independent check passing on each one. The finance lead agrees the minimum run before go-live. The licence is written into the case store with its date, the cases it was earned on, and the model and rule versions in use.

A supplier seen for the first time has no licence, so its first invoices are always reviewed, however clean they look. That one consequence removes a large share of the confident errors we used to see, because new suppliers and new layouts are where silent misreads concentrate.

A statistical argument nobody in finance wants to have becomes a rule a controller can audit. Show me the licence. Show me the cases it was earned on. Show me what would revoke it.

Evidence comes from checks that could have failed

The strongest evidence that an answer is right rarely comes from the model that produced it. It comes from checks that could have failed on their own.

For a supplier invoice they are concrete. The supplier matches the supplier master by tax number, not by name. The lines add up to the total, and the total matches the purchase order within the agreed tolerance. A goods receipt exists and covers the quantity. The proposed cost centre is valid for the entity and has been used for this supplier and material group before. Each check is deterministic, cheap and recorded against the case.

An answer that passes every check in a licensed segment goes through, whatever its score. An answer that fails one goes to a person, however high its score. The score survives only as a way to order the review queue.

Consequence overrides accuracy

A licence says how often a segment is right. It says nothing about what a mistake costs. A cost-centre error on a small invoice is fixed at month-end with a journal. A changed bank account on a supplier record can send a payment to a fraudster.

So some decisions never go straight through, whatever the licence says: changes to bank details, amounts above a limit the controller sets, suppliers on a watch list, anything touching intercompany balances near the close. The system prepares those cases fully. A person releases them.

The best segments need the closest watch

The errors that reach the ledger rarely come from messy segments. Messy suppliers fail checks and land in review, where people catch them. The damaging errors come from segments that have been reliable for months. They are licensed, reviewers no longer see them, and when the supplier changes its layout or the business reorganises cost centres, nothing in the model's self-assessment moves. Reliability removes the observation that would have caught the change.

So every licensed segment keeps a small random sample going to review permanently, and three events revoke a licence automatically: a confirmed error after straight-through processing, any change to the model or the matching rules, and a new document layout from the supplier, detected by comparing each document's structure with the layouts seen before. Revocation is not a failure. It is the system noticing that its evidence is out of date.

Licences are earned and kept on current production cases, not on the acceptance set, which is a snapshot of the past. And the controller never sees one accuracy figure for the whole system, because it is dominated by the largest suppliers and hides the segments where errors cluster. She sees the licensed segments, the cases behind each licence, and every confident error since the last report.

One question settles it

When someone shows us a system that routes on model confidence, we ask: of the cases that went straight through last month, how many were later found to be wrong, and in which segments? If nobody can answer, nobody knows how accurate the system is. They only know how sure it sounds.

Discuss Your Operation With Our Engineers.

Describe one workflow and the systems it relies on. A senior engineer responds within two business days with an initial assessment: what we would build, what we would not, and why.

A senior engineer reads every request and replies within two business days.

Or book directly: discovery call calendar