False Positive Ratio

The false positive ratio — more commonly called the false positive rate — is the proportion of legitimate cases that a control wrongly flags as suspicious. In fraud and compliance operations it is the number that determines what the control actually costs, because in almost every real system the flagged population is overwhelmingly made up of people who did nothing wrong.

It is calculated as false positives divided by all truly negative cases: of everyone who was legitimate, what share did we flag?

Also called False positive rate, false alarm rate, fall-out
Formula False positives ÷ (false positives + true negatives)
Complement Specificity — the share of legitimate cases correctly passed
Typical in sanctions screening Frequently above 90% of alerts raised
Direct cost Review labor, delayed onboarding, abandoned customers
Set by The decision threshold — not the model

Two rates that get confused

A distinction worth being precise about, because the two are routinely reported interchangeably and differ enormously when fraud is rare.

  • False positive rate asks: of all the legitimate cases, what share did we flag? It is calculated over the legitimate population.
  • False discovery rate asks: of all the cases we flagged, what share were legitimate? It is calculated over the alerts.

When someone says “our false positive rate is 95%,” they nearly always mean the second. The gap between them is the base rate, and it is where most misreadings of fraud statistics come from.

Work through it. With 1,000,000 transactions at a 0.1% fraud rate, a control catching 80% of fraud at a 1% false positive rate flags 800 real frauds and 9,990 legitimate transactions. The false positive rate is 1% — which sounds excellent — and 92.6% of the alerts are wrong. Both numbers are correct and they describe very different operational realities.

This is why an impressively low false positive rate can still produce an unworkable review queue, and why a control’s quality cannot be judged from that figure alone when the thing being detected is rare.

The trade-off that has no solution

Every threshold moves false positives and false negatives in opposite directions. Tighten to catch more fraud and more legitimate customers are caught with it; loosen to reduce friction and more fraud passes. The model determines the shape of that curve; the threshold picks a point on it. No setting escapes the trade.

What the trade costs is asymmetric, and the asymmetry is usually measured badly:

False positive False negative
What happened A legitimate case was flagged A fraudulent case passed
Who bears it The customer, and the review team The business, or the victim
Visibility Immediate — queue length, complaints Delayed, often never attributed
How it is measured Well — it generates work Poorly — undetected fraud is unlabeled
Typical bias Overweighted because it is visible Underweighted because it is not

The last row is the important one. False positives announce themselves; false negatives mostly do not. Organizations therefore tune against the error they can see, which is not necessarily the error that costs most.

Why it matters for identity verification

The framing that makes this actionable: a false positive is usually a case the system did not have enough information to resolve. Not a case it got wrong — a case it could not decide, escalated by default.

That reframes the problem. Tuning the threshold moves cases between queues without adding information, so every gain on one error is paid for on the other. Adding a discriminating attribute moves the curve itself, which is the only change that improves both at once.

Identity data is the clearest example. A sanctions screening alert on a name shared with a designated individual is unresolvable on a name alone and closes instantly against a verified date of birth. An application flagged for velocity because several submissions share a device is a different question once the applicants are verified as distinct real people. In both cases the alert disappears not because the threshold moved but because the ambiguity that created it is gone.

There is a fairness dimension worth stating. False positives are not distributed evenly. People with thin credit files, recent arrivals, those who have changed address or name, and users of older devices all deviate more from whatever the model treats as normal. Tuning against an aggregate rate hides this, and the aggregate can improve while the burden on a specific group gets worse.

What the false positive rate cannot tell you

It says nothing on its own. Quoted without the catch rate it is meaningless — a control that flags nothing has a perfect false positive rate. The pair has to be read together, at a stated operating point.

It does not capture the cost. A flagged application that resolves in ten seconds and one that takes three days and loses the customer count identically in the rate and not at all in the consequences.

It hides distribution. An aggregate rate averages across a population in which some groups are flagged far more than others.

It cannot be compared across systems with different base rates. The same control applied to populations with different fraud prevalence produces different alert compositions, which is why cross-organization benchmarking of these figures is usually not comparing anything.

Frequently asked questions

What is the difference between false positive rate and false discovery rate?

False positive rate is the share of legitimate cases that were flagged, calculated over the legitimate population. False discovery rate is the share of alerts that turned out legitimate, calculated over the alerts. When fraud is rare these diverge dramatically — a 1% false positive rate can still mean over 90% of alerts are wrong.

What is a good false positive rate?

There is no universal figure, because it is meaningless without the catch rate it was achieved at and the cost of each error for that business. The useful question is not what rate to target but what a false positive and a false negative each cost, since those two numbers determine where the threshold belongs.

Why are sanctions screening false positive rates so high?

Because matching must be fuzzy to handle transliteration and name variation, because thresholds are deliberately conservative given strict liability, and because list entries often carry too little data to discriminate on. The result is alert volumes where the large majority are name collisions rather than matches.

How do you reduce false positives without missing more fraud?

By adding information rather than moving the threshold. Threshold changes trade one error for the other along a fixed curve; a new discriminating attribute — a verified date of birth, a confirmed document number — moves the curve, which is the only way to improve both at once.

Related reading

Discover Our Solutions

Exploring our solutions is just a click away. Try our products or have a chat with one of our experts to delve deeper into what we offer.

Report
Mapping the Rise of AI-Powered Identity Fraud

AI didn't just make fraud faster. It made it a system. We analyzed millions of identity interactions to map how identity attacks are evolving across regions, attack types, and sophistication levels — and what organizations need to rethink to keep pace.

See the Data