CHAPTER 04 35 MIN READ INTERMEDIATE

SOAR and Automation

A SIEM tells you something happened. It does not, by itself, do anything about it. That gap between "alert fired" and "problem handled" is where Security Orchestration, Automation, and Response tools live. SOAR platforms connect the tools a SOC already uses, run repeatable response steps without waiting on a human for every click, and give analysts a structured record of what happened and why. Used well, automation buys analysts back the time they used to spend on repetitive lookups. Used carelessly, it can turn a single bad alert into a real outage. This chapter covers both halves of that trade-off.

SOAR playbooks automation orchestration
Building on Chapter 3: Chapter 3 covered how a SIEM collects, normalizes, and correlates log data into alerts. This chapter picks up right after that alert fires: what a SOC does with it, and how much of that "what happens next" can safely be handed to a machine.

What SOAR Adds on Top of a SIEM

Security Orchestration, Automation, and Response is really three distinct capabilities bundled under one acronym, and it helps to pull them apart before talking about the platforms that implement them.

Three capabilities, one acronym:
  • Orchestration is the connective tissue: a SOAR platform integrates with the SIEM, the EDR, the identity provider, the ticketing system, and threat intel sources, so that one tool can read from and act on all of them instead of an analyst tabbing between six consoles.
  • Automation is what runs inside that connective tissue: predefined steps that execute without a human clicking through each one, from a simple IOC lookup to a multi-stage containment sequence.
  • Response is the case management layer that ties automated and manual work together into a single record: what fired, what ran, what a human decided, and how the case closed.

None of those three capabilities is new on its own. Analysts have always pulled threat intel, opened tickets, and written up incidents. What SOAR changes is that these steps become defined, repeatable, and (where appropriate) automatic, instead of living in an analyst's head or a wiki page nobody updates. The practical effect on a SOC's day-to-day work is fewer manual, low-judgment steps and more time spent on the decisions that actually require a person.

SIEM vs. SOAR at a Glance

DimensionSIEMSOAR
Core jobCollect, normalize, and correlate log data; generate alertsAct on alerts: enrich, decide, respond, document
Primary outputAn alert, a dashboard, a query resultA completed (or partially completed) response, with a case record
Where the work happensInside log data, largely read-onlyAcross connected tools, taking write actions (isolate a host, disable an account, block an indicator)
Analyst's roleQuery, correlate, tune detection logicBuild and approve playbooks, review automated actions, handle what automation can't
Typical question answered"Did something suspicious happen?""Now that we know, what do we do about it, and can part of that be automatic?"

SOAR Doesn't Replace the SIEM

It's worth being precise about the relationship: SOAR does not replace a SIEM, and it is not a competing detection engine. It consumes the SIEM's output as its primary trigger and extends the SOC's reach past detection into the response side of the workflow. A SOC without a SOAR platform still does all three of orchestration, automation, and response; it just does them manually, through an analyst copying an IP address between four browser tabs.

Note: Not every SOC runs a dedicated SOAR product. Smaller teams sometimes build a lightweight equivalent out of scripts, scheduled tasks, and a ticketing system's webhook rules. The label matters less than the underlying discipline: predefined, repeatable, documented response steps instead of ad hoc actions re-invented on every shift.

Anatomy of a Playbook

A playbook is the unit of work inside a SOAR platform: a defined sequence of steps that runs, in whole or in part, whenever a specific condition is met. Playbooks vary enormously in complexity, from a three-step enrichment lookup to a branching workflow with a dozen decision points, but nearly all of them follow the same basic shape.

1
Trigger
An alert fires in the SIEM (or another connected tool) and matches the condition that kicks off this playbook.
→
2
Enrichment
The platform automatically gathers context: IOC reputation, user and asset details, related alerts, without an analyst asking for any of it.
→
3
Decision Point
The enriched data is evaluated against defined logic, or handed to an analyst, to determine what happens next.
→
4
Action
A response step executes, either fully automated or gated behind human approval depending on the playbook's design.
→
5
Notify / Document
The case record is updated, relevant people are notified, and the outcome is logged for later review.

Trigger Scope

The trigger stage is where scope gets defined. A playbook can be tied to a single alert rule, a category of alerts, or a broader condition such as "any alert tagged credential access above a certain severity." Overly broad triggers cause playbooks to run against alerts they were never really designed for, so trigger conditions deserve the same care as the detection logic that fires them.

Enrichment: The Easy Win

Enrichment is usually the highest-value, lowest-risk stage, which is why it's almost always the first thing SOCs automate. Pulling an IOC's reputation score, checking whether a user account has flagged activity elsewhere, or looking up whether an asset is a Tier 0 server takes an analyst a few minutes of manual work per alert; a playbook does it in seconds, every time, without forgetting a step.

The Decision Point

The decision point is where a playbook's design intent shows up most clearly. Some decision points are pure logic: if the reputation score crosses a threshold, branch one way, otherwise branch another. Others explicitly route to a human: presenting the enriched context to an analyst and waiting for a judgment call before anything happens. Which pattern is correct depends heavily on what the downstream action is, which is the subject of the next section.

Note: A playbook that only automates enrichment and stops at the decision point, handing everything off to an analyst from there, is still a meaningful automation win. Full end-to-end automation through the action stage is not the only useful outcome; it's simply the more advanced one.

Orchestration vs. Automation, and Human-in-the-Loop vs. Full Automation

These terms get used loosely, often interchangeably, and the imprecision causes real confusion when a SOC is deciding what to build. It's worth separating two distinct axes: what the platform is doing (orchestration vs. automation) and who is allowed to pull the trigger on the final action (human-in-the-loop vs. full automation).

TermWhat It MeansExample
OrchestrationConnecting and coordinating multiple tools so they can be acted on from one placeA playbook can query the EDR, check the identity provider, and open a ticket, all from the same workflow
AutomationExecuting a defined action without a human performing it manually, step by stepAutomatically pulling an IOC's reputation score the moment an alert containing that IOC fires
Human-in-the-loopAutomation runs up to a point, then pauses for a person to approve or reject the next stepA playbook enriches an alert and drafts a recommended action, but an analyst must click approve before an account gets disabled
Full automationThe entire sequence, including the final response action, runs without waiting for human approvalA playbook automatically blocks a known-malicious IP at the firewall as soon as reputation intel confirms it

Orchestration Is Not Automation

Orchestration and automation are not the same thing, even though most SOAR marketing treats them as a package deal. A platform can orchestrate (connect to five tools) without automating much of anything, if every step still requires a human to click "run." Conversely, a single automated action, like an enrichment lookup, doesn't require much orchestration at all if it only touches one external source. Most real SOAR deployments do both at once, but keeping the distinction clear helps when troubleshooting: "the integration is broken" (orchestration) is a different problem from "the playbook logic made the wrong call" (automation).

Where to Draw the Approval Line

The more consequential decision is where to draw the human-in-the-loop line, and that decision should track directly against two properties of the action: how confident the playbook can be that the action is correct, and how easily the action can be reversed if it turns out to be wrong.

Low-risk, high-confidence, reversible actions are good candidates for full automation. An enrichment lookup doesn't change anything in the environment; it only gathers information, so even a wrong or irrelevant result costs a few seconds of wasted context, not an outage. Adding an IOC to a watchlist, tagging a case with additional metadata, or opening a ticket are similarly low-stakes: easy to undo, and the downside of a mistake is small.

High-impact, hard-to-reverse actions belong behind a human approval gate. Disabling a user account, isolating a host from the network, or blocking a wide network range all have real operational cost if triggered on a false positive: a legitimate user locked out mid-task, a production server pulled offline, a supplier's traffic silently dropped. These actions are exactly the ones where a thirty-second human sanity check, "does this match what I'd expect from a real incident," prevents a much larger cleanup effort later.

Tip: A useful rule of thumb when deciding whether a playbook step should be automated end-to-end: ask what it costs the business if the trigger turns out to be a false positive and the action runs anyway. If the answer is "nothing meaningful," automate it. If the answer is "a person has to explain this to their manager," gate it behind approval.

Common Enrichment Patterns

Enrichment is where most SOAR programs deliver their earliest and most consistent return, because the actions involved are read-only, fast, and directly reduce the manual work an L1 analyst would otherwise repeat on every alert. A few patterns show up across nearly every SOC that has built out automation, regardless of which platform they run.

PatternWhat It Does
IOC reputation lookupsThe moment an alert contains an IP address, domain, file hash, or URL, a playbook can automatically check that indicator against threat intelligence sources and attach the result to the case before an analyst even opens it. Instead of an analyst manually pasting an indicator into three different lookup tools, the context is already sitting in the alert when they start working it.
User and asset context pullsAn alert that names a username or a hostname is more useful once it's paired with who that user is (department, manager, whether they're currently traveling or on leave, recent HR flags if the SOC has visibility into that) and what that asset is (server role, criticality tier, patch status, normal login pattern). Pulling this context automatically turns a bare username into something an analyst can actually triage against, rather than a string they have to go look up manually.
Related-alert correlationA single alert rarely tells the whole story. A playbook can automatically query the SIEM for other alerts involving the same user, host, or indicator within a defined time window, surfacing a pattern (five failed logins followed by a success, followed by unusual data access) that would otherwise require an analyst to notice the connection themselves across separate alert tickets.

What these patterns have in common is that they don't change anything in the environment. They gather and assemble information that already exists, just faster and more consistently than a person clicking through consoles. That's precisely why enrichment is the natural first automation target for a SOC building out a SOAR program: the upside is real and the downside of a bad enrichment result is close to zero.

Note: Enrichment automation is most valuable when its output is easy for an analyst to scan quickly. A playbook that dumps ten paragraphs of raw API responses into a case note isn't actually saving anyone time; it's just relocating the reading burden. Good enrichment design summarizes and highlights what changed the analyst's picture of the alert, not everything the integration was capable of returning.

The Risk of Over-Automation

Automation doesn't just speed up correct decisions. It speeds up wrong ones too, and it does so consistently, at machine speed, without the natural hesitation a person feels before taking an action they're not fully sure about. A single low-confidence alert that would have taken an analyst a few minutes to reason through and probably dismiss can, if it's wired into a fully automated response chain, trigger a real action against production infrastructure before anyone has looked at it.

Warning: An automated playbook that isolates a host based on a single alert with no correlation or human review can turn a false positive into an actual outage. If that host happens to be a critical server, the SOC has now caused the exact kind of business disruption it exists to prevent, and it did so faster and with less oversight than a human ever would have.

Blast Radius

The core discipline here is thinking in terms of blast radius: not just "is this action reversible," but "how much of the business does this action touch if it fires on bad data." Blocking a single low-reputation IP has a small blast radius. Disabling an account used by one person has a moderate one. Isolating a shared file server or a domain controller has a large one, because the fallout extends to everyone who depends on that system, not just the account or host the alert named.

The same logic that governs where to put a human approval gate applies here and should scale with blast radius: the wider the potential fallout, the more scrutiny an action deserves before it runs, and the more conservative the confidence threshold should be before automation is trusted to pull that trigger alone.

Staged Rollout

The practical way SOCs manage this risk is a staged rollout of automation, rather than flipping a playbook straight to full automatic on day one.

1
Notify-Only Mode
A new playbook runs its enrichment and decision logic, but instead of taking the response action, it simply tells an analyst what it would have done. This lets the team observe how the playbook behaves against real alert volume, catch logic errors or over-broad triggers, and build confidence in its accuracy without any risk of a wrong action actually executing.
→
2
Approval-Gated Automation
Once a playbook has proven reliable in notify-only mode, it can graduate here: the enrichment and decision steps still run automatically, but the final action waits for a human to click approve, as described in the human-in-the-loop pattern earlier in this chapter.
→
3
Full Automation
Only after a playbook has a track record of correct recommendations at the approval-gated stage, ideally across a meaningful volume of real alerts and enough time to catch edge cases, does it make sense to move it here, for the narrow set of actions that are genuinely low-risk and high-confidence.

Skipping stages is the most common cause of an automation program losing the SOC's trust. One bad outcome from a playbook that went straight to full automation, especially one with a wide blast radius, tends to set the whole automation initiative back further than a slower, staged rollout ever would have cost in lost time.

Common SOAR Platforms

Several established products implement the orchestration, automation, and response capabilities covered in this chapter. This isn't an endorsement or a comparison of feature depth, just a factual grounding in the names a SOC analyst is likely to encounter on the job or in an interview.

PlatformVendorNotes
Cortex XSOARPalo Alto NetworksSOAR platform supporting playbook automation, case management, and third-party integrations across security tooling
Splunk SOARSplunk (part of Cisco)SOAR platform built to work alongside Splunk's SIEM offerings, with playbook-based automation and orchestration
Microsoft Sentinel PlaybooksMicrosoftAutomation and response capability built into Microsoft Sentinel, implemented on top of Azure Logic Apps

Beyond the platform-specific interface, the underlying concepts transfer directly: a trigger, an enrichment stage, a decision point, an action, and a documentation step, gated by whatever mix of human approval and full automation the SOC has decided each action deserves. Learning to reason about that structure matters more than memorizing any one product's menu layout, since the specific tool a SOC uses will vary by employer far more than the underlying playbook logic will.

Key Takeaways

  • SOAR is three capabilities bundled together: orchestration (connecting tools), automation (running actions without a human), and response (case management and documentation). A SIEM detects; SOAR acts on what it detects.
  • A playbook's typical shape is trigger, enrichment, decision point, action, notify/document. Enrichment is usually the safest and highest-value stage to automate first.
  • Orchestration (connecting tools) and automation (running actions unattended) are related but distinct. Human-in-the-loop and full automation describe a separate axis: whether a person approves the final action.
  • Low-risk, high-confidence, reversible actions like enrichment lookups are good candidates for full automation. High-impact, hard-to-reverse actions like disabling an account or isolating a host deserve a human approval gate.
  • Common enrichment patterns include IOC reputation lookups, user and asset context pulls, and related-alert correlation, none of which change anything in the environment.
  • Over-automation risk comes down to blast radius: a fully automated action on a single low-confidence alert can turn a false positive into a real business disruption. Staged rollout (notify-only, then approval-gated, then full auto) manages that risk.
  • Cortex XSOAR, Splunk SOAR, and Microsoft Sentinel Playbooks (built on Azure Logic Apps) are common products implementing these concepts.

Knowledge Check

Click an answer to reveal the explanation.

Which best describes the difference between what a SIEM does and what a SOAR platform does?

Correct answer: B. A SIEM's job is detection: turning raw log data into alerts through correlation. A SOAR platform picks up from there, consuming those alerts as its primary trigger and orchestrating the response across connected tools. SOAR doesn't replace SIEM detection capability (C), and neither tool is limited to a single data type (D).

A SOC is deciding whether a playbook step should run as full automation or wait for human approval. Which factor pair should drive that decision?

Correct answer: B. The human-in-the-loop decision should track confidence and reversibility: low-risk, high-confidence, easily reversible actions (like enrichment lookups) suit full automation, while high-impact, hard-to-reverse actions (like disabling an account or isolating a host) warrant an approval gate. Platform tenure (A), operating system (C), and shift staffing (D) aren't the relevant factors.

A SOC just built a new playbook that automatically isolates any host flagged by a single medium-confidence alert, with no human review step. What is the most likely risk with this design?

Correct answer: B. This is exactly the over-automation risk covered in the chapter: a hard-to-reverse, potentially high-blast-radius action (host isolation) wired to full automation on a single alert with no correlation or human review can amplify a false positive into a genuine business disruption. Staged rollout (notify-only, then approval-gated, then full auto) exists precisely to catch this before it happens in production.