Workshop Resources

Respond · Reference sheet

Incident Kickoff Toolkit

The first hour of an incident: the moment it's declared through the first status brief. Use it to build the incident-response section of your DR/BC plan — or hand it out as-is.

In a live incident, keep the quick reference card at hand instead. This document is for building your organization's own procedures ahead of time.

1. Incident Start Checklist

Run this once, at the moment an incident is declared. The goal is a fast, consistent kickoff — not a perfect one.

1.1 Is it incident-shaped?

Check the situation against the severity matrix before you open an incident. If it's clearly a P4, it may just be a ticket.

LevelCriteriaResponse
P1 — CriticalBusiness outage, all users affected.Immediate fix required.
P2 — HighMajor impact, multiple users/services down.Fast resolution needed.
P3 — MediumLimited impact, workaround possible.Handled in normal SLA.
P4 — LowMinor issue, cosmetic or single user.Planned resolution.

Not sure which level? Default to P3 and re-triage once you know more. It's easier to stand a response down than to catch up from behind.

1.2 Apply STOP before acting

1.3 The checklist

1.4 Roles in incident response

Based on the Incident Command System (ICS). Assign only what the incident needs.

RoleOwns
Incident CommanderOwns the response. Makes the call others can't or shouldn't. One person, full authority, for the duration.
Public Information OfficerOwns external and stakeholder-facing communication. Nobody else talks to customers or press.
Safety OfficerWatches the responders — sleep, stress, physical and operational safety. Can pause the response.
Liaison OfficerCoordinates with outside parties: vendors, partners, other internal teams not directly responding.
Operations Section ChiefRuns the technical response — the people actually fixing the problem.
Planning Section ChiefTracks status, next steps, and the parking lot of things to deal with later.
Logistics Section ChiefGets responders what they need — access, tools, food, rest.
Finance Section ChiefTracks cost and contractual impact — SLAs, vendor spend, credits owed.

Decision-making structure: authority runs through the Incident Commander. Section Chiefs run their sections; they escalate decisions outside their authority to the IC, not around them.

2. Briefing Checklist

A brief keeps leadership, stakeholders, and responders aligned without pulling responders off the fix. It is a sync, not a status meeting — keep it short and keep it moving.

2.1 Who's in the room

Incident Commander, active Section Chiefs, the Public Information Officer, and one stakeholder representative. Anyone actively fixing the problem should not be in the brief — send someone in their place, or get the update from the Run Sheet.

2.2 Cadence

Set the default cadence by priority level, and confirm the next brief time at the end of every brief. Adjust to what your organization can sustain.

PrioritySuggested cadence
P1 — CriticalEvery 30 minutes
P2 — HighHourly
P3 — MediumOnce per shift
P4 — LowAs needed

2.3 Agenda — same order, every time

2.4 Ground rules

Stick to facts. Blame slows the response down and makes people less likely to surface bad news early — which is when it's most useful.

2.5 Stakeholder update (separate from the internal brief)

Shorter, plainer, and less technical than the internal brief. Owned by the Public Information Officer. Three lines, no more:

Draft this before the first external update goes out, and reuse the same three-line structure for every update after — consistency reads as competence.

3. Run Sheet Template

A single chronological record of what happened, when, and who did it. It's the incident's source of truth: it feeds the stakeholder updates, the post-mortem, and any required documentation.

3.1 How to use it

3.2 Fields

FieldWhat goes here
Time24-hour clock, to the minute.
Action / EventWhat happened or what was done — one clear sentence.
OwnerWho did it or is doing it.
StatusLogged / In progress / Done / Blocked.
NotesAnything needed for context later — links, error codes, decisions made.

3.3 Example

TimeAction / EventOwnerStatusNotes
09:14Alert fired — payments API error rate > 20%.J. AlvarezLoggedAuto-detected via monitoring.
09:17Incident declared, P1. IC assigned.J. AlvarezDoneIC: J. Alvarez
09:22Ops began rollback of latest deploy.R. OseiIn progressETA 10 min

3.4 Template — for use

TimeAction / EventOwnerStatusNotes
     
     
     
     
     
     
     
     
     
     

After the incident: this log is the primary input to the post-mortem (what happened, how we found out, how we responded, how we recovered, what we're doing about it) and to any required incident documentation.