Free DevOps maturity audit for new clientsBook a 30-min call

Reliability & Performance

SRE, On-Call & Incident Response

Error budgets, runbooks and a page rotation your engineers can live with.

The work

What this actually does

Reliability becomes manageable when it is defined. We help you name the few journeys that matter, agree what good service looks like in numbers, and set an error budget that decides how fast you can ship. That turns reliability from an argument about feelings into a policy the whole team can read.

Day to day, the work is incident response. We review the incidents you have had, tighten the escalation path, write runbooks for the pages that recur, and run a blameless postmortem process that produces dated actions with owners rather than a document nobody opens again. The measure of success is fewer pages and shorter ones.

If any of these sound familiar
  • Every incident turns into the same arguments afterwards
  • We cannot say whether last month was better or worse
  • The same failures happen again with no recorded fix
  • One engineer carries the whole on-call burden

Scope

What's included

Every engagement on this page covers the following, sized to your setup rather than delivered as a fixed package. If something here is not relevant to you, it comes off the scope and off the price.

Service-level indicators

We define good and bad events for each critical journey — successful checkout, fast search — so reliability can be measured rather than described.

Error budget policy

A written agreement on what happens when the budget runs out: which releases pause, who decides, and how the decision is reversed once reliability recovers.

Runbooks and playbooks

Each repeating alert gets a numbered procedure with the exact commands, the expected result, and the point at which the responder should escalate instead of persisting.

Blameless postmortems

We facilitate the review, capture the timeline from real telemetry, and turn findings into owned actions tracked in the same backlog as feature work.

On-call rotation design

Schedules are sized to the real page volume, with a secondary responder, documented swaps and a rule that nobody is on call alone forever.

Reliability reporting

A monthly view of incidents, budget burn and open actions, written for both engineers and the people who fund the platform.

What changes

What teams typically see

15 minMedian incident response
1 pageError budget policy
monthlyReliability review cadence

Handover

What you keep

Everything produced during the engagement is yours: the repositories, the accounts, the documentation. There is no proprietary layer and nothing to unlicense if you take the work in-house.

  • Written service-level indicators and objectives per journey
  • Error budget policy signed off by engineering
  • Runbook template and a first set of runbooks
  • Postmortem template with tracked action items
  • On-call rotation plan with escalation responsibilities

Tooling

Tools we use here

A starting point, not a requirement. We work in whatever you already run wherever it does the job.

Prometheus
Grafana
PagerDuty
Datadog
Kubernetes
awsAWS

How it runs

From first call to handover

The same four steps on every engagement. You see each one before it starts and can stop at any of them.

  1. 01

    Agree what reliability means

    We interview the teams and the business to find the journeys that actually matter, then propose indicators that can be measured from data you already have.

  2. 02

    Set budgets and policy

    Targets are agreed in writing, including the consequences of breaching them, because a budget with no policy behind it changes nothing.

  3. 03

    Rebuild the incident path

    Escalation, communication and handover are documented and rehearsed, so an incident has a known shape rather than being improvised each time.

  4. 04

    Run the review cycle

    We chair the first postmortems and the first monthly reliability review, then hand the format to your team once the habit has formed.

Questions

Asked before we start

Is an error budget just a way to stop releases?

It is a way to make the trade-off explicit rather than accidental. When the budget is healthy you ship faster with less ceremony; when it is spent, the policy says what pauses. Most teams find it removes arguments rather than causing them.

Do small teams need formal SRE practice?

Smaller teams get the most from it, because there is less slack to absorb a bad night. You can run the whole thing on a page of markdown and one review a month; the ceremony is optional, the clarity is not.

How does this fit with existing ITIL processes?

They can coexist. We map incident severities and change records onto the process you already run, so the new practice produces the evidence your change board needs instead of competing with it.

What if we do not have the telemetry to measure SLOs?

Then that is the first piece of work. We start with what your load balancer and application already record, which is often enough for availability and latency, and add instrumentation only where a gap genuinely blocks a decision.

Reliability & Performance

Often needed alongside this

Reliability & Performance

Alerting & On-Call Tooling

Fewer, better alerts that route to the right person — and stay quiet when nothing is actually wrong.

  • PagerDuty, Opsgenie and Grafana OnCall
  • Alert tuning to kill chronic noise
  • Escalation policies and rotation schedules
See the full service
Reliability & Performance

Monitoring & Observability

Know what broke, why it broke and who it affects — before your customers have to tell you.

  • Prometheus, Grafana, Datadog & CloudWatch
  • OpenTelemetry tracing and structured logging
  • SLO dashboards and actionable alert routing
See the full service
Optimisation & Advisory

24/7 Managed DevOps Support

An on-call team on the other end of the pager, with agreed response targets and monthly incident reporting.

  • 24/7 monitoring, triage and incident response
  • Contractual response and resolution targets
  • Monthly reliability and cost reporting
See the full service

Worth a conversation about SRE, On-Call & Incident Response?

Bring the specific problem. We will tell you honestly whether this is the service that fixes it, and what it would take.