Alert rule review
We go through every existing rule, ask what action the responder should take, and either rewrite it around a clear signal or delete it.
Reliability & Performance
Fewer pages, sent to the right person, and silence when nothing is wrong.
The work
An on-call rota is only sustainable when the pager means something. Teams usually arrive with hundreds of rules, most of them firing on symptoms nobody can act on, and a rotation that has learned to ignore the noise. We rebuild the alert set around user-visible conditions, then make sure each one reaches a named person who can do something about it.
The plumbing matters as much as the rules: schedules, escalation policies, notification channels and quiet hours all need to reflect how your team actually works, including holidays and handovers. We configure PagerDuty, Opsgenie or Grafana OnCall and test it with a real page during the working day, so the first time it fires is not at three in the morning.
Scope
Every engagement on this page covers the following, sized to your setup rather than delivered as a fixed package. If something here is not relevant to you, it comes off the scope and off the price.
We go through every existing rule, ask what action the responder should take, and either rewrite it around a clear signal or delete it.
Alerts are mapped to the owning team with escalation paths, so an unanswered page moves on after a defined interval instead of sitting unread over a weekend.
Grouping, inhibition and dependency rules collapse a cascade of related failures into one page, which is usually the single largest reduction in volume.
A written routine covers what the outgoing engineer hands over, how schedule swaps work, and where the current state of open incidents is recorded.
After each shift we log which pages were useful and which were not, then adjust the rules, so the rota keeps improving rather than calcifying.
Synthetic probes from outside your network test the journeys customers actually use, catching failures that internal metrics miss entirely.
What changes
Handover
Everything produced during the engagement is yours: the repositories, the accounts, the documentation. There is no proprietary layer and nothing to unlicense if you take the work in-house.
Tooling
A starting point, not a requirement. We work in whatever you already run wherever it does the job.
How it runs
The same four steps on every engagement. You see each one before it starts and can stop at any of them.
We pull the last month of pages and rank them by volume and by how often they were resolved with no action at all.
Rules are rebuilt around user-visible conditions and clear thresholds, with the noisy infrastructure checks demoted to dashboards where they belong.
Schedules, escalation steps and notification preferences are set up and written down, including who is accountable when a page is not acknowledged.
We trigger each critical alert during working hours, watch it arrive on every channel, and fix the gaps while everyone is awake.
Questions
That is the whole point, and it needs judgement rather than a blanket threshold. We look at what each page was supposed to detect, keep coverage for genuine failure, and remove only the rules that have demonstrably never led to action.
No. We work with PagerDuty, Opsgenie, Grafana OnCall or whatever you run today. The vendor is rarely the problem; unowned alerts and missing escalation paths are, and those travel with you either way.
That disagreement is useful evidence. We default to a written owner and an explicit action for each rule, and if neither can be produced, the alert is removed. Removing a page is reversible; burnout is not.
We separate what genuinely cannot wait until morning from what can, and only the former pages anyone. Everything else waits in a queue with the same owner and severity, reviewed at the start of the next working day.
Reliability & Performance
Error budgets, runbooks and an escalation path that keeps 3 a.m. pages rare and short.
Know what broke, why it broke and who it affects — before your customers have to tell you.
Diagrams, decision records and step-by-step runbooks that turn a 3 a.m. page into a checklist.
Bring the specific problem. We will tell you honestly whether this is the service that fixes it, and what it would take.