Grouping a flood of errors into decisions
Thousands of events become one issue with an owner.
- The surface
- An error arriving from production: how it is grouped, what the issue page shows, and how it gets assigned and resolved.
- What the user wants
- I want to know which of today's errors is actually hurting users, and fix that one.
Grouping into issues
Incoming events are fingerprinted, usually from the stack trace and error type, and identical failures collapse into a single issue with an event count and a count of affected users.
Aggregate before you alert
A raw error stream is unreadable during an incident, because one broken code path can generate an enormous volume of near-identical lines. Collapsing them turns noise into a list of distinct problems. Counting affected users separately from events is the difference between one person in a retry loop and a broad outage.
The stack trace with context
The issue page shows the stack trace with source context, and with source maps or debug symbols configured it resolves to the original code rather than the built output.
Restore the developer's own frame of reference
A minified trace is technically complete and practically useless, so it pushes the engineer into a local reproduction before they can even start. Mapping back to real filenames and lines lets them recognise the code. Recognition is the whole game in triage.
Breadcrumbs
Each event carries a timeline of what happened before the failure, such as navigation, network calls and console output.
Reconstruct the path, not just the endpoint
Most bugs are only reproducible with the sequence that caused them, and the user who hit it will never be able to describe that sequence. A recorded trail replaces the bug report nobody writes well. It converts an unreproducible ticket into an investigable one.
Release and suspect commit
Issues are tagged with the release they appeared in, and with commit data connected, Sentry surfaces likely responsible changes and their author.
Attach the cause to the symptom
The most useful question during triage is what changed, and answering it manually means correlating deploy times against error onset by hand. Linking the two collapses that work to a glance. Naming an author routes the issue without a meeting.
Resolve in next release
An issue can be marked resolved in the next release, and if the same error reappears after that release ships, it reopens and is flagged as a regression.
State that verifies itself
Resolved usually means a human clicked a button, which is a claim rather than a fact. Tying resolution to a release lets the data confirm or reject the claim automatically. Regression detection is what stops the issue list from filling with optimistic closures.
Alert rules over aggregates
Alerts can be configured on conditions such as a new issue appearing or an issue exceeding a threshold over a window, rather than firing for every event.
Alert on the change, not the level
An alert per event is an alert nobody reads, and after a week the channel is muted. Firing on novelty or on a rate change keeps the signal proportional to the decision being requested. The threshold is a product decision about what the team will actually stop for.
From a flood to one decision
What not to copy
- Grouping is heuristic, and both failure modes hurt. Over-grouping hides a distinct bug inside a busy issue, and under-grouping splits one bug into a dozen entries that each look minor.
- Event-based pricing pushes teams to sample, and sampling decisions get made under cost pressure rather than diagnostic need. The rare error you most want to see is the one most likely to be dropped.
- The issue list grows faster than any team triages it. Without a discipline nobody schedules, it becomes a backlog people scroll past, and the tool has no answer for that.
- Rich context means production data ends up in a third-party system. Breadcrumbs and local variables are exactly where personal data leaks, and scrubbing is configuration the team has to get right in advance.
The takeaway
Before adding an alert, ask what decision it triggers and whether the aggregate would answer it better.
Finished the teardown? Bank it and the day counts toward your run.
Where the principles come from
- Site Reliability Engineering (alerting philosophy), Google
- The Design of Everyday Things (error and feedback), Don Norman
- Error Message Guidelines, NN/g
Written from public behaviour of the product, not from inside it. Interfaces change often, so treat the flow described here as of the time of writing and check the live product before quoting it.
