Base rate neglect
Given a vivid specific case, people forget how rare the category is.
When combining a prior probability with new evidence, people lean on the evidence and under-weight the prior. The consequence is well known in diagnostics: a highly accurate test for a rare condition still produces mostly false positives, and most people badly overestimate the chance that a positive result is real. The neglect is strongest when the evidence is individuating and concrete.
How it shows up in software
Any product that flags, scores, or alerts runs into this. A fraud alert, a security warning, a health screening result, and a moderation flag all carry a false positive rate that users do not intuit. If the flagged condition is rare, most flags are wrong, and a user who does not see the prior will over-trust the alert and then learn to ignore it when the false positives pile up.
Using it well
- State results as natural frequencies in the UI: 'about 3 of every 100 accounts flagged this way turn out to be fraudulent'.
- Show the base rate beside the score in any review queue, so the reviewer anchors on prevalence and not only on the match.
- Tune alert thresholds against prevalence, not accuracy alone, since a rare event makes a high-accuracy detector noisy.
- Track whether users still act on your alerts after a month. Falling action rates mean the false positive rate is teaching them to ignore you.
Where it turns manipulative
- Publishing a model's accuracy without its precision on a rare class is the standard way to oversell a classifier, and it lands on users as unwarranted certainty.
- Reporting a relative risk increase without the absolute base rate makes a tiny effect sound large, a habit that is well documented in health and security marketing.
- Framing a flagged account or a screening result as a verdict rather than a signal, when the prior makes most flags wrong.
Where you have seen it
Stripe Radar
Payments get a risk score with a stated risk level and the rule evidence that produced it, not a binary verdict alone.
23andMe
Genetic health reports present the user's result alongside population prevalence and state that the report is not diagnostic.
Google Safe Browsing
Interstitial warnings describe what was detected and offer a details path rather than only blocking.
What the research says
- Kahneman and Tversky, 1973Well evidenced
Given a personality sketch and told it was drawn from a pool that was 70 percent lawyers, participants judged the profession from the sketch and largely ignored the stated proportion.
- Casscells, Schoenberger and Graboys, 1978Well evidenced
Asked about a disease with a 1 in 1000 prevalence and a test with a 5 percent false positive rate, most staff and students at a Harvard teaching hospital answered around 95 percent. The correct answer is roughly 2 percent.
- Gigerenzer and Hoffrage, 1995Well evidenced
Restating the same problems as natural frequencies (10 out of 1000) rather than probabilities raised correct Bayesian answers dramatically across participants and problems.
This is the most actionable finding here. Base rate neglect is heavily a presentation problem, so the fix is to change the format before concluding that your users cannot reason.
Grades are a judgement about the evidence, not about how useful the idea is. Plenty of contested effects are still worth knowing, as long as you do not cite them as settled.
