Dashboards that must survive being woken by them
Tags, template variables, and a monitor that pages a human.
- The surface
- Building a service dashboard, filtering it with template variables, and configuring a monitor that notifies a team.
- What the user wants
- I want to know whether the service is healthy right now, and I want to be told before a customer tells me.
Tags as the query language
Metrics arrive with tags such as service, environment and region, and queries filter and group by those tags rather than by metric name alone.
Model the dimensions, not the charts
A naming scheme that encodes meaning into the metric string forces a new metric for every combination anyone might want. Tags let one metric answer questions nobody anticipated at instrumentation time. The quality of every later dashboard is decided by the tagging discipline set early.
Template variables
A dashboard exposes dropdowns bound to tags, so one dashboard serves every service or environment instead of being duplicated per team.
Parameterise instead of duplicating
Copied dashboards diverge quietly, and during an incident nobody knows which copy is current. One parameterised dashboard keeps a single definition of what healthy looks like. It also means an improvement made by one team lands for everyone.
Layout as a reading order
Dashboards are grids of widgets where placement is chosen by the author, and effective ones put user-facing symptoms at the top with causes below.
Visual hierarchy under stress
A dashboard is read by a tired person at three in the morning, and that is the only use case worth designing for. Putting error rate and latency above resource graphs means the first glance answers are users affected. A grid with no hierarchy makes everything equally urgent, which is the same as nothing being urgent.
Monitors with states
A monitor evaluates a query against thresholds over a window and holds a state of OK, warn, alert or no data, with recovery thresholds to prevent flapping.
Hysteresis in the alerting loop
A single threshold on a noisy signal produces a stream of alerts and recoveries that trains the team to ignore the channel. Separate trigger and recovery thresholds give the state somewhere to settle. A distinct no data state matters because silence and health look identical otherwise.
The notification body
Monitor messages are templated and can carry the triggering values, a graph snapshot, tags, and a link to a runbook written by the team.
Put the first three minutes of response into the page
The person woken by an alert starts from nothing, and every second of context gathering happens while the incident continues. Embedding the values, the graph and the runbook link collapses orientation into reading one message. It also enforces the discipline of writing a runbook while you configure the alert.
Downtimes and muting
Scheduled downtimes suppress monitors during known maintenance, scoped by tag rather than by silencing everything.
A sanctioned way to be quiet
If there is no legitimate way to suppress an alert, engineers will find an illegitimate one and it will outlive the reason for it. Scoped and scheduled suppression keeps the exception visible and bounded. The alternative is a permanently muted channel nobody remembers muting.
Which monitors deserve to wake somebody
What not to copy
- Pricing spans hosts, custom metrics, indexed logs and spans, and each has its own cardinality trap. Teams discover the cost of a tag after they have shipped it everywhere, and the fix is an instrumentation change.
- The tool makes it trivial to create dashboards and offers nothing for retiring them. Large organisations end up with hundreds, most stale, and finding the right one during an incident becomes its own problem.
- Onboarding assumes an admin has already installed agents, set up integrations and defined the tagging scheme. A new engineer arriving after that work sees a product of enormous surface area with no visible entry point.
- Observability is a category with deep switching costs. Dashboards, monitors and years of tagging conventions do not migrate, and that immobility is a real part of the pricing power.
The takeaway
Write the runbook link into the alert template, or accept that the runbook does not exist.
Finished the teardown? Bank it and the day counts toward your run.
Where the principles come from
- Site Reliability Engineering (monitoring and alerting), Google
- The Visual Display of Quantitative Information, Edward Tufte
- The Design of Everyday Things (signifiers and feedback), Don Norman
Written from public behaviour of the product, not from inside it. Interfaces change often, so treat the flow described here as of the time of writing and check the live product before quoting it.
