CHAPTER 04 30 MIN READ INTERMEDIATE

AI in the SOC: Detection and Triage

A SOC analyst opening their queue on a Tuesday morning is not short on alerts, they are short on time to look at each one properly. AI is now embedded in that workflow at three distinct points: behind the alert (anomaly detection deciding what fires at all), on top of the alert (copilots enriching and summarizing what came in), and beside the analyst (natural-language interfaces that translate a question into a query). This chapter walks through each layer honestly, including where they help and where they quietly introduce new risk.

AI-assisted triageSOC copilotsanomaly detectionUEBA
Before you start: this chapter assumes you already know what a SOC analyst's triage queue looks like day to day, and that you have some familiarity with signature-based detection (a rule that fires on a known pattern). It does not assume prior exposure to machine learning concepts, those are introduced as needed.

Anomaly detection and UEBA: unsupervised ML in practice

User and Entity Behavior Analytics (UEBA) is the most mature deployment of machine learning inside the SOC. Instead of writing a rule that says "alert when X happens," a model learns what "normal" looks like for a specific user, host, or entity, then flags moments where behavior deviates from that baseline by a statistically significant margin.

The baseline is built per entity, not globally. A service account querying a database every sixty seconds all day is normal for that account, even though the same volume from a marketing laptop would be wildly abnormal. This is unsupervised learning: nobody hand-labels "malicious" or "fine," the model infers unusual entirely from the shape of the entity's own history and its peer group.

What a UEBA baseline typically tracks:
  • Login time-of-day and geographic pattern for a given user
  • Typical volume and destination of data a host or account moves
  • Which systems, shares, or applications an entity routinely touches versus never touches
  • Peer-group comparison, how this entity's behavior compares to others with a similar role
  • Rate and sequence of actions, not just whether an action happened but how fast and in what order

The payoff: this catches behavior nobody wrote a rule for. A compromised account used in a technically permitted way, wrong hours, an unusual resource, an oversized data pull, will not trip a single signature, because no individual action is inherently malicious. It is the deviation from that entity's own history that is the signal, the "unknown unknown" case: a technique nobody has documented, expressed through behavior that doesn't look like the account it's coming from.

The tradeoff is just as real. Unsupervised anomaly detection tends to run a higher false-positive rate than a well-tuned signature, for a few structural reasons:

ReasonWhy it inflates false positives
Unusual is not the same as maliciousA legitimate but rare event, an employee's first trip to a new office, a one-time bulk export for a quarterly audit, looks exactly like an anomaly to the model because it is one. It just isn't an attack.
The baselining period is a weak spotA model that has only observed two weeks of an entity's behavior has a thin, noisy picture of "normal," and over-flags during that window until enough history accumulates.
High behavioral variance is harder to baselineA consulting firm where analysts touch different client systems every week is inherently harder to baseline than a stable back-office environment, and the false-positive rate reflects that regardless of tuning quality.

The detection engineering module covers false-positive tuning for signature-based rules in depth. The same underlying discipline, understanding what "normal" looks like in your environment before deciding what should alert, applies here too, just expressed through baseline tuning, sensitivity thresholds, and peer-group configuration instead of rule logic.

AI-assisted alert triage and enrichment

Once an alert exists, whether from a signature or a UEBA anomaly, it still has to be worked. This is where LLMs and ML classifiers have made the most visible inroads into day-to-day SOC operations: pre-triage of the incoming queue, a layer sitting between the raw alert and the analyst.

1
Alert fires
Signature or anomaly-based detection generates a raw alert with minimal context attached.
→
2
Auto-enrichment
System pulls process tree, host/user history, and threat intel context without analyst action.
→
3
Plain-language summary
LLM condenses the enriched data into a short narrative and a suggested severity.
→
4
Draft disposition
A recommended close or escalate decision is proposed, not applied.
→
5
Analyst confirms
A human reviews the summary and evidence, then accepts, edits, or overrides the recommendation.

It is worth being precise about what this is and is not. This is augmentation of Tier 1 triage work, not a replacement for it. The reason the line is drawn there rather than further along is not tradition, it is risk math: an LLM can hallucinate, stating a process is signed when it is not, misattributing an IOC to the wrong campaign, or simply misjudging context it was never given. The cost of an incorrectly auto-closed true positive is high, a real intrusion dismissed as noise can sit undetected for weeks, while the cost of a human spending thirty extra seconds confirming a correct AI suggestion is negligible by comparison. That asymmetry is why human review of AI-suggested dispositions remains standard practice rather than an optional safeguard, even at organizations that have otherwise gone all-in on AI tooling.

SOC copilots and natural-language interfaces

A second, related category of tooling is now common inside SIEM and XDR platforms: chat-style copilots that let an analyst ask a question in plain English (e.g. "show me all failed logins from this host in the last 24 hours") and get back a query, its results, and a summary, without writing the query language themselves.

What it looks like
BenefitLowers the barrier to a query language that can take months to get comfortable with, and speeds up routine lookups an experienced analyst could write fast but a newer one would have to look up syntax for.
RiskStructural, not incidental: an analyst who can't read the underlying query has no independent way to verify the copilot translated their intent correctly. An ambiguous request ("failed logins") could mean interactive logons only, or every auth attempt including service accounts and RDP, and a narrower interpretation the analyst never catches produces a summary that looks complete while quietly missing part of the picture.

Chapter 5 goes deeper into natural-language-to-query workflows in the hunting and rule-authoring context, including how experienced practitioners use these tools without losing the underlying skill they're built on top of.

Where AI-assisted triage breaks down

Worth being direct about the failure modes rather than treating them as footnotes, they're the reason every one of these tools is deployed with a human still in the loop.

Failure modeWhat it looks like
HallucinationA model states something incorrect with the same confident tone as something correct, no built-in signal distinguishes them. E.g. mischaracterizing a process as a known Windows binary when it's actually a lookalike, or misattributing an IOC to the wrong campaign.
Missing organizational contextNot really the model's fault: it can't judge context it was never given. Unusual admin activity at 2 a.m. during a scheduled maintenance window looks identical to a real intrusion unless the maintenance calendar was explicitly fed into its context.
Automation bias (complacency)Trust in AI-suggested dispositions builds with repeated correct outcomes, leading to less scrutiny over time, a well-documented human factors pattern. The tool being usually right is exactly what makes its rare wrong call dangerous.

These are why human-in-the-loop review stays standard for consequential dispositions (closing an alert benign, escalating to IR) rather than a transitional step toward full automation. They're structural properties of a probabilistic summarizer lacking full organizational context, not bugs the next model version patches out.

Practical guidance for adopting AI triage tools

None of the above is an argument against using these tools, it is an argument for adopting them deliberately. A SOC evaluating or rolling out AI-assisted triage should treat it the same way it would treat any other change to a process that carries operational risk.

ConsiderationWhy it matters
Pilot against ground truth before trusting it at scaleRun the tool's suggestions alongside human dispositions for a defined period and compare them directly, rather than assuming vendor-claimed accuracy transfers to your environment.
Keep a human decision point for closes and escalationsAnything that ends an alert's lifecycle, benign close or IR escalation, should require a human action to actually execute, not just a suggestion an analyst can ignore by doing nothing.
Log AI-suggested dispositions separately from human onesAuditing the two streams independently is the only way to measure drift, bias, or a slow decline in suggestion quality over time, rather than discovering it after an incident.
Don't let copilots replace learning the query languageAnalysts who never learn the underlying syntax cannot verify what a copilot ran or catch a misinterpreted request, a gap covered in more depth in chapter 2 of the Detection Engineering module.

None of these considerations require slowing adoption to a crawl. They require building in the checkpoints that let a SOC find out it is wrong before that wrongness compounds into a missed intrusion or a burned-out team that has quietly stopped reading what the AI hands them.

Key Takeaways

  • UEBA and unsupervised anomaly detection baseline normal behavior per entity and flag statistically significant deviations, catching novel behavior signatures miss, at the cost of a higher false-positive rate, especially early in baselining or in high-variance environments.
  • AI-assisted triage tools summarize, enrich, and propose a disposition for incoming alerts, augmenting Tier 1 work rather than replacing it, because the cost of an incorrectly auto-closed true positive is too high to automate away entirely.
  • SOC copilots translate natural-language questions into queries, lowering the barrier to routine lookups, but an analyst who cannot read the underlying query has no way to verify the translation was correct.
  • Hallucination, missing organizational context, and automation bias are structural limitations, not temporary bugs, which is why human-in-the-loop review stays standard for consequential dispositions.
  • Adopting these tools responsibly means piloting against ground truth, keeping a human checkpoint on closes and escalations, auditing AI dispositions separately, and not letting copilots substitute for learning the query language underneath them.

Knowledge Check

Click an answer to reveal the explanation.

Compared to a well-tuned signature-based detection, what is the core tradeoff of unsupervised anomaly detection like UEBA?

Unsupervised anomaly detection learns a per-entity baseline of normal behavior and flags statistically significant deviations from it, which means it can surface "unknown unknown" behavior that no signature covers. The same mechanism that makes it useful for novel detection also makes it prone to flagging rare-but-legitimate events, particularly while the baseline is still thin or in environments where behavior naturally varies a lot.

Why does human review of AI-suggested alert dispositions remain standard practice rather than being replaced by full automation?

The reasoning is a risk asymmetry: confirming a correct AI suggestion costs a human a few seconds, while an incorrectly auto-closed true positive can let a real intrusion go undetected. Since models can state incorrect things confidently and cannot see context they were never given, that asymmetry is what keeps a human decision point in place for consequential dispositions.

An analyst has accepted a SOC copilot's "likely benign" recommendation correctly for months, and now reviews each new recommendation with less scrutiny than they used to, assuming the tool is probably right again. What is this pattern called?

This is automation bias, also called complacency: trust in an AI's suggestions builds up with repeated correct outcomes, leading to less scrutiny over time, even though the tool being usually right is exactly what makes its rare wrong call more likely to slip through unnoticed.