A threshold is a policy, not a number
A threshold is a policy, not a number Somewhere in a payments codebase there is a line that says: approve automatically when confidence is above 0.8. Nobody remembers the afternoon it was written. The number has three likely origins and all of them are bad — it was the first value that made the demo behave, it was copied from a vendor example, or it was chosen because 0.8 sounds strict without…
In the realm of payments codebases, there exists a crucial line that states: approve automatically when confidence is above 0.8. However, the genesis of this number remains shrouded in mystery, likely stemming from one of three unrelated sources. The significance of this value cannot be overstated, as it determines whether refunds are executed automatically by a machine or require human intervention.
Setting this threshold improperly can have far-reaching consequences, with easy-to-ignore refunds being processed automatically, while the majority of cases are funneled into a human queue that was not designed to handle the increased workload. The underlying system is ordinary, with refund requests processed through arithmetic, a decision model, and human review for the more challenging cases.
The core issue lies in the single number that dictates which path a given case takes, and how this decision is made based on taste rather than empirical measurement. The rule layer, which appears straightforward and trustworthy, fails silently, allowing the business to move forward without noticing the underlying flaw. Meanwhile, the decision model operates differently, providing a distribution rather than a binary outcome.
When faced with a case outside its training data, it returns a number without raising an error, leading to a failure mode that goes unnoticed. The third destination is not a more sophisticated model, but rather a different question altogether - which of two viable options should be chosen. This decision is not apparent until the ticket ages and a decision is made that was never made in the first place, resembling a backlog rather than a mistake.
The failure modes are not the sole focus of this article, as each one carries a price tag. The gate's threshold is a policy, not a number, and the decision on who bears the cost lies with the user. The 0.8 value originates from various sources, none of which provide a clear indication of its true meaning. Accuracy measures the overall correctness of the system, while confidence provides an estimate of the model's hit rate.
These two concepts answer different questions and should not be conflated. Calibration is essential to bridge the gap between reported confidence and actual error rates. Choosing the threshold is not a simple math problem, but rather a matter of consequences that the model has not been exposed to. Setting the threshold too low can result in irreversible damage, such as refunds being processed without proper review.
On the other hand, setting it too high may lead to a backlog of cases, causing delays and hidden costs. Both directions have their drawbacks, and the ideal threshold depends on the specific consequences at hand. The reliability report serves as the only artifact that quantifies the system's performance, breaking down the calls by confidence bands and their accuracy rates.
This data is published for transparency, allowing users to make informed decisions based on the available evidence. Ultimately, the choice of threshold is a policy decision, not a technical one, and there is no one-size-fits-all solution. The curve must be considered before a threshold can be set, ensuring that the decision aligns with the organization's goals and tolerances.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.