H3AD-REF / PLAYBOOKS / RANSOMWARE IR

Ransomware IR.
The First Hour Decides The Next Month.

A static process reference: what to do in the first 15 minutes, the IR lifecycle applied specifically to ransomware, containment and eradication checklists, and the mistakes that turn a contained incident into a repeat one. Not a substitute for an org's own IR plan or legal counsel, just the reference. If this started as a single execution alert, see the Malware / Execution Alert playbook for the triage steps that came before this one.

Immediate Triage

The first 15 minutes set the ceiling on how much damage is recoverable. Every action here is about buying time without destroying evidence.

First 15 Minutes, In Order

1. Isolate, don't shut down
├── Pull the network cable or disable the NIC on affected hosts
└── Do not power off or reboot — encryption keys and unencrypted data can live in memory,
    and a reboot can trigger a second-stage payload some ransomware families carry
    // isolation stops lateral spread without destroying volatile evidence a reboot would lose

2. Identify scope before assuming it's contained
├── Check file shares, backup servers, and domain controllers, not just the first host that alerted
└── Ransomware that's been staged for days often has already reached backup infrastructure
    // scope is almost always wider than the first alert suggests

3. Preserve evidence before remediating
├── Take a memory capture and disk image of at least one representative affected host, if tooling allows
└── Remediating before imaging destroys the ability to identify patient zero or the initial vector
    // this step gets skipped under pressure more than any other, and it's the one that's hardest to undo

4. Activate the IR plan and notify stakeholders
├── Legal, executive leadership, and (if applicable) cyber insurance get notified now, not after triage
└── Insurance policies frequently require notification within a specific window to preserve coverage
    // waiting to "have more information first" is the most common reason coverage gets disputed later

The IR Lifecycle, Applied To Ransomware

The standard six-phase incident response lifecycle, with what each phase actually means when the incident is ransomware specifically.

PHASE 1

Preparation

Everything done before the incident, that determines how the rest of this goes
Offline/immutable backups tested for restore, not just existence. An IR retainer or plan that names who does what. EDR coverage on servers, not just endpoints, since ransomware frequently pivots through under-monitored infrastructure first.
PHASE 2

Identification

Confirming what happened, not just that something did
Identify the ransomware family if possible (ransom note format, file extension, known TTPs), the initial access vector, and how long the actor had access before deploying the payload (T1486). The deployment is usually the last step of a much longer intrusion.
PHASE 3

Containment

Stopping the spread without destroying what's needed to eradicate it
Network segmentation, disabling compromised accounts, and isolating backup infrastructure specifically, since backups are a common secondary target. See the dedicated containment checklist below.
PHASE 4

Eradication

Removing the actor's access, not just the visible payload
Rotating every credential the actor could have touched, removing persistence mechanisms, and patching the initial access vector. Eradicating only the ransomware binary while leaving the access that delivered it is how re-encryption incidents happen within weeks.
PHASE 5

Recovery

Restoring operations from a position of confirmed clean state
Restore from backups verified to predate the compromise, not just the most recent backup. Rebuild from known-clean images where restoration confidence is low. Bring systems back online in priority order, not all at once.
PHASE 6

Lessons Learned

The phase most often skipped once the pressure is off
A post-incident review that identifies the initial access vector's root cause (unpatched service, phished credential, exposed RDP) and closes it, not just a summary of what happened. Skipping this is why the same organization gets hit twice.

Containment Checklist

The specific actions that stop spread, in the order they matter most.

CONTAIN

Isolate Backup Infrastructure First

Backups are a deliberate target, not collateral damage
Modern ransomware operators specifically hunt for and encrypt or delete backup infrastructure (T1490) before deploying the main payload. Verify backup systems are isolated and backup integrity is confirmed before doing anything else network-wide.
CONTAIN

Rotate Credentials By Priority

Domain admin and service accounts first, not alphabetically
Prioritize accounts with the broadest access: domain admins, service accounts with elevated rights, and anything the compromised host had cached credentials for. A full password reset across the org can wait; the highest-value accounts can't.
CONTAIN

Segment Before You Sweep

Contain the blast radius before trying to characterize its size
Isolate affected network segments (VLANs, firewall rules) before running org-wide sweeps for indicators. Scanning an uncontained network while the actor still has access can tip them off that they've been detected.
CONTAIN

Decide On The Payment Question Early, Formally

This is a legal and executive decision, never a technical one made under pressure
Whether to engage with the ransom demand is a decision for legal counsel, executives, and (if applicable) law enforcement and cyber insurance, made deliberately, not by whoever is on the incident bridge at 2am.

Eradication & Recovery

Getting back online without bringing the actor's access back online with it.

RECOVER

Identify Patient Zero

The initial access point, not the first host to show symptoms
The host that alerted first is rarely the host the actor entered through. Timeline analysis across authentication logs, EDR telemetry, and network logs is usually required to trace back to the actual entry point.
RECOVER

Sweep For Persistence Before Restoring

Scheduled tasks, new services, and web shells outlive the initial payload
Ransomware operators frequently establish persistence mechanisms well before deploying the encryptor. Restoring a system without checking for these leaves the door open for a second incident using the same access.
RECOVER

Restore Vs Rebuild, Per System

Not a single org-wide decision
Restore from a verified-clean backup where confidence is high. Rebuild from a known-clean image where the compromise timeline is unclear or the system was directly used by the actor. Mixing both approaches across different systems is normal, not indecisive.
RECOVER

Bring Systems Back In Priority Order

Business-critical first, but verified-clean always comes before critical
A priority list drafted during preparation (not improvised during recovery) prevents both unnecessary downtime and the temptation to restore an unverified system early because it's important.

Common Mistakes

The decisions made under pressure that turn a contained incident into a longer, worse one.

MISTAKE

Rebooting Affected Systems

Instinctive, and frequently destructive
A reboot can destroy volatile evidence needed to identify the actor's access method, and some ransomware families use a reboot as a trigger for a second-stage action. Isolate; don't power-cycle.
MISTAKE

Restoring From An Infected Backup

Re-encrypts the environment on the very system meant to save it
If the actor had access before the last backup ran, that backup can already be compromised. Verify a backup predates the intrusion timeline, not just the ransomware deployment, before trusting it for recovery.
MISTAKE

Paying Without Legal And Law Enforcement Involvement

A decision with legal and regulatory weight, not just a technical one
Payment can carry sanctions exposure depending on the threat actor's identity, and most cyber insurance policies have specific notification and process requirements that void coverage if skipped.
MISTAKE

Skipping Forensic Imaging Before Remediation

The fastest fix and the biggest information loss
Wiping and reimaging a host immediately restores service but destroys the ability to determine the initial access vector, meaning the root cause never gets fixed and the same vector stays open.

Escalation Decision Points

Who needs to be in the room, and when, before a technical decision becomes a legal or reputational one.

ESCALATE

Legal Counsel

Before any communication with the threat actor, and before any public statement
Regulatory notification obligations (breach disclosure laws vary by jurisdiction and data type) and the payment decision both sit with legal, not IT or security.
ESCALATE

Law Enforcement

Earlier than most organizations default to
Agencies like the FBI (in the US) or a national equivalent elsewhere often have threat-actor-specific intelligence, including decryption keys recovered from prior takedowns, that isn't available anywhere else.
ESCALATE

Cyber Insurance

Within the notification window the policy specifies, not "once things settle down"
Insurers often mandate use of specific IR firms or approved vendors as a condition of coverage. Contacting them late, or engaging outside vendors first, can affect what's reimbursable.
ESCALATE

Executive Leadership And PR

Before, not after, information starts leaking externally
Ransomware groups increasingly run their own leak sites and contact customers or press directly to pressure payment. Leadership and communications need to know before finding out from a journalist or a customer.