Responding to an Incident
Pager, the on-call engineer at Incident Command, has one rule on the wall: "Users first, cause second." When production breaks, the natural reaction is to open the code and start hunting for the bug. That feels productive, but every minute spent debugging while users are failing is a minute of damage you could have stopped. A calm routine beats clever debugging under pressure.
Step 1: Confirm the impact
An alert is a claim, not a fact. Before you act, spend a minute or two checking:
- Are users actually affected? Look at the error rate, latency and traffic dashboards, not just the alert.
- Who and what? All users or one region, every page or only checkout?
- Since when? The exact start time is your best clue later.
Write the answers down in the incident channel as you go. "Checkout errors at 30% since 14:05, other pages fine" tells everyone, including you, what you are dealing with. If the alert turns out to be noise, say so and fix the alert afterwards.
Step 2: Stop the bleeding
Your first goal is to make users stop failing, even before you understand why. The fastest levers are usually:
- Roll back the most recent deploy if the problem started around then. It is quick, well-tested and easy to undo.
- Turn off a feature flag for the new feature that looks involved.
- Scale up or restart if the service is overloaded and that buys time.
- Fail gracefully: hide a broken widget or disable an optional feature so the core flow works.
Rolling back before finding the bug feels like giving up. It isn't. The logs, metrics and the bad commit are still there after the rollback, and you can find the cause in peace while users are fine.
Step 3: Communicate
Other people need to know what is happening: support is answering angry emails, your manager is being asked questions, and users see errors. Post short updates at a steady rhythm:
14:12 Checkout is failing for about 30% of users since 14:05.
We are rolling back this afternoon's deploy. Next update by 14:30.
Say what is affected, what you are doing, and when the next update comes. Don't guess at causes in public. On a bigger incident, one person fixes and another communicates, so the fixer isn't interrupted every two minutes.
Step 4: Find the cause
Once users are safe, investigate. Start with what changed, because most incidents follow a change:
- deploys of your app, including the ones you think were harmless
- configuration or environment variable changes
- database migrations
- a dependency or third-party service that started failing
- a traffic spike from a launch, a newsletter or a bot
Line up the start time from Step 1 with these events. Then use the logs: filter the failing requests by error message and endpoint, pick one request id, and read its whole story. Form one guess at a time and check it against the data, instead of changing several things at once.
Step 5: A blameless postmortem
After the incident, write a short postmortem: timeline, impact, cause, what went well, what didn't, and follow-ups. Blameless means you ask "how did our system allow this?" rather than "who did this?". If one engineer ran a migration that locked a big table, the useful question is why nothing warned them, not whose fault it was. People who fear blame hide details, and without details you can't fix anything.
Follow-ups must be concrete, with an owner and a date:
- Weak: "Be more careful with migrations."
- Strong: "Add
lock_timeoutto the migration runner and a CI check for missing indexes. Owner: Sam, by March 14."
A postmortem that ends with "be more careful" guarantees the same incident again.
Resources
Curated resources for this node are on the way. Use what you already know how to search for, and check back soon.