Architecture diagrams
Context, container and dependency diagrams generated from live infrastructure where possible, then reviewed with the engineers who own each service so errors get corrected early.
Optimisation & Advisory
Diagrams, decision records and runbooks that turn a 3 a.m. page into a checklist.
The work
Documentation fails when it drifts. A diagram is drawn during a design review, the service gains three dependencies, and six months later the picture is a polite fiction. We rebuild the core set from what is actually running: the cloud APIs, the cluster, the pipeline definitions and the alert rules. Then we put it somewhere your team will maintain it, next to the code.
The runbooks matter most. Every alert that can wake someone up gets a page-long procedure: what the alert means, how to confirm it, the commands to run, and the point at which to stop and escalate. We test them during business hours by following them literally, because a runbook nobody has executed is a document, not a control.
Scope
Every engagement on this page covers the following, sized to your setup rather than delivered as a fixed package. If something here is not relevant to you, it comes off the scope and off the price.
Context, container and dependency diagrams generated from live infrastructure where possible, then reviewed with the engineers who own each service so errors get corrected early.
One runbook per alert, covering symptoms, confirmation steps, remediation commands, rollback and the escalation trigger, stored beside the code that produces the alert.
Short architecture decision records capturing the options considered, the choice made and the consequences accepted, so a question from two years ago has an answer.
A first-week guide for each role: accounts to request, repositories to read, a small first task and the people to ask about each system.
Everything lives in Markdown in your repositories, reviewed in pull requests and published by the same pipeline that ships the service it describes.
Scheduled checks that flag runbooks untouched after a related service change and diagrams that no longer match the infrastructure they claim to show.
What changes
Handover
Everything produced during the engagement is yours: the repositories, the accounts, the documentation. There is no proprietary layer and nothing to unlicense if you take the work in-house.
Tooling
A starting point, not a requirement. We work in whatever you already run wherever it does the job.
How it runs
The same four steps on every engagement. You see each one before it starts and can stop at any of them.
We collect the diagrams, wikis, runbooks and tribal knowledge that exist today, then compare them against the systems and alerts that are actually live.
Architecture diagrams, alert runbooks and decision records are drafted with your engineers, not for them, so the content is accurate and someone already owns it.
We execute each procedure in a non-production environment, correct the steps that do not work, and add the rollback and escalation detail that was missing.
Documentation lands in your repositories with a review checklist, a stale-content check and an owner named against every page so drift is visible.
Questions
Because it lives in the repositories they already work in, changes with the code in the same pull request, and is checked automatically when it goes stale. If a page needs a separate tool and a separate login, it will rot; ours does not.
That is a common starting point and part of the work. We read the code, the infrastructure and the logs, interview whoever has been closest to it, and mark our uncertainty explicitly rather than guessing. Gaps get listed as questions for the team to confirm.
Your team does, from day one. Everything sits in your repositories under your accounts, with named owners per area and a review cadence. We are not the only people who can edit it, and that is deliberate.
Where a diagram can be generated from infrastructure, we generate it and review the output on a schedule. Where it cannot, the diagram is tied to a service owner and checked in the same review as the next change to that service.
Optimisation & Advisory
A maturity assessment that produces a prioritised roadmap, then hands the team the skills to run it.
Your engineers learn to run the platform themselves — through pairing, workshops and internal runbooks.
Know what broke, why it broke and who it affects — before your customers have to tell you.
Bring the specific problem. We will tell you honestly whether this is the service that fixes it, and what it would take.