Skip to content
All posts
Incident Response7 min read

Your first hour: an incident plan that survives being tired

Plans fail at three in the morning for predictable reasons. What to strip out of yours, what to write down, and how to rehearse it so the first hour is recall rather than improvisation.

Daniel FreijeIncident Response Lead

I have been on enough three in the morning calls to have formed a strong opinion about incident response plans: almost all of them are too long, and the length is not a neutral property. A forty page plan is not a more thorough version of a four page plan. It is a plan that will not be opened.

The first hour of an incident is where the avoidable damage happens, and it is also the hour when everybody involved is least capable. People are half awake, missing context, and under a pressure that makes reading comprehension noticeably worse. Design for that person, not for the auditor.

What the first hour has to produce

Not resolution. The first hour has four jobs, and a plan that helps with these four is worth more than one that documents everything.

  1. Somebody is in charge, by name, and everybody on the call knows who that is.
  2. The blast radius has a first estimate, even a rough one, written where others can see it.
  3. A decision has been made about containment, including the decision to wait and watch.
  4. A timeline has been started, with timestamps, because reconstructing it later is far harder than keeping it now.

What to strip out

Most plans carry material that belongs somewhere else. Move it there and the plan becomes usable.

  • Regulatory notification text. Necessary, but not in hour one. Put it in an appendix with a clear deadline against it.
  • Definitions and scope statements written for a framework. Keep them in the policy document the auditor reads.
  • Long decision trees. Under stress people follow the first branch and stop. Two or three clear questions beat a flowchart.
  • Tool-specific instructions that go stale. Link to the runbook instead, and keep the runbook next to the tool.

What to write down instead

Our template fits on a page and a half. It contains the escalation path with real phone numbers, the severity definitions in one sentence each, the names of the people who are allowed to authorise disruptive containment, the out-of-band communication channel, and a short list of the questions the commander asks in the first ten minutes.

That last item is the piece most plans lack. Write the questions down. What do we know, how do we know it, what would change our estimate, who else needs to be awake, and what is the worst case if we are wrong. Five questions, asked out loud, in that order, will structure the first hour better than any flowchart.

Rehearsal is the part that cannot be skipped

A plan that has never been exercised is a hypothesis. The exercise does not need to be elaborate: ninety minutes, the real on-call people, no advance warning about the scenario, and somebody taking notes on where the plan was ambiguous.

The same three failures surface almost every time. Nobody knows who is allowed to take production offline. The communication channel depends on the system that is currently compromised. And the person with the credentials required for containment is not on the escalation list.

All three are cheap to fix once you have seen them. None of them will be visible from reading the document. Run the exercise, fix what it exposes, and the plan stops being paperwork and starts being the thing that shortens your worst night of the year.

Working on this

If this describes something you are dealing with, a scoping call costs nothing and takes about half an hour. We will tell you honestly whether it needs an engagement or an afternoon of your own team's time.

Request a Security Assessment
  • Incident Response

    Logging you will actually use when it matters

    Most teams collect too much of the wrong thing and too little of the right thing. A practical review of what to keep, how long to keep it, and how to check your pipeline before an incident tests it for you.

    Priya Raghunathan7 min read
    Read post
  • Threat Research

    The authorisation bug you will not find with a scanner

    Broken object-level authorisation is the finding we report most often, and almost none of it is caught by automated tooling. Here is how we test for it by hand, and the three patterns that keep producing it.

    Tomas Erlend9 min read
    Read post

Next step

Turn a worry into a scoped engagement

Tell us what prompted you to read this. We will tell you what we would test first, and what it would cost.

Expires in

Limited time offer

We rebuilt your site for you. Claim it and we handle everything transfer, hosting, and your domain. Then update it anytime, just by asking AI.

Host for only$8 per monthBilled yearly
Claim limited offer now