Overconfidence effect
Confidence outruns accuracy, and confidence intervals come out far too narrow.
Overconfidence is three distinct things that get conflated. Overestimation is thinking you performed better than you did. Overplacement is thinking you are better than others. Overprecision is holding beliefs with more certainty than the evidence supports, and it is the most robust of the three. They do not move together and can point in opposite directions on the same task.
How it shows up in software
Users overestimate what they can do without help, which is why documentation goes unread and advanced features stay undiscovered. Teams overestimate how well they understand their users, which is why five interviews get treated as settled truth. Forecasts and model outputs carry the same problem: a ranked list with no uncertainty reads as more certain than the underlying data supports.
Using it well
- Put uncertainty in the interface where the output is uncertain: confidence bands, ranges, and 'not enough data' states.
- Teach through the flow rather than through docs, since a user who is confident will not go looking for instructions.
- In forecasting, ask for an 80 percent range and then track how often the truth landed inside it. Calibration improves with feedback, slowly.
- Give experiments a pre-registered sample size: How many observations a test needs before its result means anything. and decision rule, so a confident reading cannot stop the test early.
Where it turns manipulative
- Shipping an AI feature that states answers in flat, confident prose regardless of underlying certainty transfers your model's overprecision to the user.
- Certification and score UI that implies mastery from a short assessment sells confidence rather than competence.
- Presenting a single-point forecast to a customer because a range looks weaker, when the range is what you actually know.
Where you have seen it
Google Maps
Arrival times are given with traffic-based ranges on some routes rather than a single fixed time.
FiveThirtyEight forecasts
Election models were published as probability distributions with visible intervals rather than a single predicted outcome.
GitHub Copilot
Suggestions appear as greyed inline text requiring explicit acceptance rather than being inserted directly.
What the research says
- Moore and Healy, 2008Well evidenced
Separated the three forms and showed people overestimate on hard tasks while underestimating on easy ones, and overplace on easy tasks while underplacing on hard ones, which reconciles a lot of conflicting earlier results.
- Lichtenstein and Fischhoff, 1977Well evidenced
Across general knowledge questions, items answered with stated 100 percent certainty were wrong a meaningful share of the time, and 90 percent confidence intervals captured the true value far less than 90 percent of the time.
- Erev, Wallsten and Budescu, 1994Mixed evidence
Showed that the same data can display over- or underconfidence depending on how it is aggregated, because of regression effects, so some measured overconfidence is a statistical artifact.
Overprecision survives this critique best and is the form worth designing around. Broad claims that people are simply overconfident are too loose to act on.
Grades are a judgement about the evidence, not about how useful the idea is. Plenty of contested effects are still worth knowing, as long as you do not cite them as settled.
