A runbook is a written, step-by-step procedure for carrying out a routine IT operation or responding to a known problem. It might explain how to restart a failed service, fail over a database, rotate a certificate, onboard a new site or restore a server from backup. The point is that someone other than the original expert can follow it and get the same result, whether that’s a new hire, an on-call engineer at 3 a.m. or an outside provider. Runbooks can be documents, entries in a knowledge base or, increasingly, automated workflows.
At a glance
- A runbook documents how to do one specific task or handle one known issue, in order, with checks along the way.
- It reduces reliance on individual memory and makes results more consistent.
- Common uses include routine maintenance, incident response, failover and recovery, and onboarding.
- Many runbook steps can be automated; steps that need judgment usually stay manual.
- Runbooks need regular review and testing, or they drift out of date as systems change.
What problem it solves
In many IT teams, critical knowledge lives in a few people’s heads. When the person who knows how to fix the billing server is on vacation, an outage lasts longer. When an engineer leaves, the team rediscovers procedures by trial and error. When a network operations center (NOC) or MSP takes over monitoring, it can only act on what is written down.
Runbooks capture that knowledge in a form others can use. Common tasks are done more consistently. On-call staff can respond to known issues without waking up the one expert. Recovery steps are clear before they’re needed, which can shorten mean time to recovery (MTTR). And when you change providers or staff, the procedures stay with the organization.
How it works
What a runbook contains. A good runbook typically includes:
- Purpose and trigger: what it’s for and when to use it, such as a specific alert.
- Prerequisites: access, tools and approvals needed.
- Steps: numbered actions in order, with expected results.
- Verification: how to confirm the task worked.
- Failure handling: what to do if a step fails, and when to stop.
- Escalation: who to contact if it doesn’t resolve the issue, often referencing an escalation matrix.
- Ownership and review date: who maintains it and when it was last checked.
Types. Operational runbooks cover routine tasks like patching, backups and user provisioning. Incident runbooks cover known failure modes, such as a full disk or a failed circuit. Recovery runbooks cover larger events and often sit within business continuity and disaster recovery (BCDR) plans.
Manual, semi-automated or automated. A runbook may be a document a person follows, a script that runs some steps, or a fully automated workflow triggered by an alert. Security teams often automate response procedures in security orchestration, automation and response (SOAR) platforms, and DevOps teams build them into deployment and monitoring tools.
Testing and upkeep. Runbooks are tested during maintenance windows, failover tests and tabletop exercises, and updated after incidents reveal gaps.
When it matters for buyers
- When onboarding an MSP or NOC. Providers act on documented procedures; agree who writes and maintains them.
- When key staff leave. Documenting procedures before someone departs avoids losing knowledge.
- After an outage that took too long. Missing or outdated runbooks are a frequent finding in post-incident reviews.
- When preparing for audits or insurance. Documented recovery and incident response (IR) procedures are commonly requested.
- When planning automation. Well-written runbooks are the starting point for automating routine work.
Our incident response overview covers how documented procedures fit into preparing for security incidents.
Questions to ask vendors
- Do you maintain runbooks for our environment, and can we access them?
- Who writes and approves runbooks, and how often are they reviewed?
- Which procedures are automated, and which require a person?
- What actions will you take on your own under a runbook, and which require our approval?
- How do you test recovery runbooks, and can we see results?
- If we end the contract, will we receive all runbooks and related documentation?
How it differs from a playbook and SOAR
The terms overlap and are used inconsistently. Generally, a runbook is specific and procedural: the exact steps for one task or issue. A playbook is broader and situational: how the organization responds to a type of event, such as a ransomware attack, including roles, decisions and communications, and it may call several runbooks. In security, “playbook” often refers to an automated workflow inside a SOAR platform, which is effectively an automated runbook for security alerts. A runbook can exist in any form, from a page in a wiki to a script, while SOAR is a specific category of security tool that executes such procedures.
