What are your first 3 steps in an incident? #1
Hive80-lab
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
A question for ops folks here: what are the first three things you actually do when an incident starts?
I've been reading public postmortems (Cloudflare, GitLab, Fastly) and the pattern is blunt: the teams that come out okay aren't the ones with the biggest runbooks — they're the ones with the clearest first 30 minutes. Under stress you don't rise to your wiki; you fall to your checklist.
The 3-step opener that keeps showing up in every good postmortem:
After that: handoffs written down as you go (not reconstructed at the end), and an out-of-band comms channel — vendor dashboards are frequently behind the thing that's down.
For transparency: I run a small ops-tools shop, and I've packaged this as a free one-page checklist (no signup): https://hive80lab.gumroad.com/l/first-30-minutes — steal it and adapt it. If you want the fuller kit (incident log templates, escalation ladder, postmortem skeleton), that's here: https://hive80lab.gumroad.com/l/ops-starter-kit
Curious what your first 3 steps are — especially the unwritten ones people only learn from a bad night.
All reactions