AI in the SOC: Detection and Triage
A SOC analyst opening their queue on a Tuesday morning is not short on alerts, they are short on time to look at each one properly. AI is now embedded in that workflow at three distinct points: behind the alert (anomaly detection deciding what fires at all), on top of the alert (copilots enriching and summarizing what came in), and beside the analyst (natural-language interfaces that translate a question into a query). This chapter walks through each layer honestly, including where they help and where they quietly introduce new risk.
Anomaly detection and UEBA: unsupervised ML in practice
User and Entity Behavior Analytics (UEBA) is the most mature deployment of machine learning inside the SOC. Instead of writing a rule that says "alert when X happens," a model learns what "normal" looks like for a specific user, host, or entity, then flags moments where behavior deviates from that baseline by a statistically significant margin.
The baseline is built per entity, not globally. A service account querying a database every sixty seconds all day is normal for that account, even though the same volume from a marketing laptop would be wildly abnormal. This is unsupervised learning: nobody hand-labels "malicious" or "fine," the model infers unusual entirely from the shape of the entity's own history and its peer group.
- Login time-of-day and geographic pattern for a given user
- Typical volume and destination of data a host or account moves
- Which systems, shares, or applications an entity routinely touches versus never touches
- Peer-group comparison, how this entity's behavior compares to others with a similar role
- Rate and sequence of actions, not just whether an action happened but how fast and in what order
The payoff: this catches behavior nobody wrote a rule for. A compromised account used in a technically permitted way, wrong hours, an unusual resource, an oversized data pull, will not trip a single signature, because no individual action is inherently malicious. It is the deviation from that entity's own history that is the signal, the "unknown unknown" case: a technique nobody has documented, expressed through behavior that doesn't look like the account it's coming from.
The tradeoff is just as real. Unsupervised anomaly detection tends to run a higher false-positive rate than a well-tuned signature, for a few structural reasons:
| Reason | Why it inflates false positives |
|---|---|
| Unusual is not the same as malicious | A legitimate but rare event, an employee's first trip to a new office, a one-time bulk export for a quarterly audit, looks exactly like an anomaly to the model because it is one. It just isn't an attack. |
| The baselining period is a weak spot | A model that has only observed two weeks of an entity's behavior has a thin, noisy picture of "normal," and over-flags during that window until enough history accumulates. |
| High behavioral variance is harder to baseline | A consulting firm where analysts touch different client systems every week is inherently harder to baseline than a stable back-office environment, and the false-positive rate reflects that regardless of tuning quality. |
The detection engineering module covers false-positive tuning for signature-based rules in depth. The same underlying discipline, understanding what "normal" looks like in your environment before deciding what should alert, applies here too, just expressed through baseline tuning, sensitivity thresholds, and peer-group configuration instead of rule logic.
AI-assisted alert triage and enrichment
Once an alert exists, whether from a signature or a UEBA anomaly, it still has to be worked. This is where LLMs and ML classifiers have made the most visible inroads into day-to-day SOC operations: pre-triage of the incoming queue, a layer sitting between the raw alert and the analyst.
It is worth being precise about what this is and is not. This is augmentation of Tier 1 triage work, not a replacement for it. The reason the line is drawn there rather than further along is not tradition, it is risk math: an LLM can hallucinate, stating a process is signed when it is not, misattributing an IOC to the wrong campaign, or simply misjudging context it was never given. The cost of an incorrectly auto-closed true positive is high, a real intrusion dismissed as noise can sit undetected for weeks, while the cost of a human spending thirty extra seconds confirming a correct AI suggestion is negligible by comparison. That asymmetry is why human review of AI-suggested dispositions remains standard practice rather than an optional safeguard, even at organizations that have otherwise gone all-in on AI tooling.
SOC copilots and natural-language interfaces
A second, related category of tooling is now common inside SIEM and XDR platforms: chat-style copilots that let an analyst ask a question in plain English (e.g. "show me all failed logins from this host in the last 24 hours") and get back a query, its results, and a summary, without writing the query language themselves.
| What it looks like | |
|---|---|
| Benefit | Lowers the barrier to a query language that can take months to get comfortable with, and speeds up routine lookups an experienced analyst could write fast but a newer one would have to look up syntax for. |
| Risk | Structural, not incidental: an analyst who can't read the underlying query has no independent way to verify the copilot translated their intent correctly. An ambiguous request ("failed logins") could mean interactive logons only, or every auth attempt including service accounts and RDP, and a narrower interpretation the analyst never catches produces a summary that looks complete while quietly missing part of the picture. |
Chapter 5 goes deeper into natural-language-to-query workflows in the hunting and rule-authoring context, including how experienced practitioners use these tools without losing the underlying skill they're built on top of.
Where AI-assisted triage breaks down
Worth being direct about the failure modes rather than treating them as footnotes, they're the reason every one of these tools is deployed with a human still in the loop.
| Failure mode | What it looks like |
|---|---|
| Hallucination | A model states something incorrect with the same confident tone as something correct, no built-in signal distinguishes them. E.g. mischaracterizing a process as a known Windows binary when it's actually a lookalike, or misattributing an IOC to the wrong campaign. |
| Missing organizational context | Not really the model's fault: it can't judge context it was never given. Unusual admin activity at 2 a.m. during a scheduled maintenance window looks identical to a real intrusion unless the maintenance calendar was explicitly fed into its context. |
| Automation bias (complacency) | Trust in AI-suggested dispositions builds with repeated correct outcomes, leading to less scrutiny over time, a well-documented human factors pattern. The tool being usually right is exactly what makes its rare wrong call dangerous. |
These are why human-in-the-loop review stays standard for consequential dispositions (closing an alert benign, escalating to IR) rather than a transitional step toward full automation. They're structural properties of a probabilistic summarizer lacking full organizational context, not bugs the next model version patches out.
Practical guidance for adopting AI triage tools
None of the above is an argument against using these tools, it is an argument for adopting them deliberately. A SOC evaluating or rolling out AI-assisted triage should treat it the same way it would treat any other change to a process that carries operational risk.
| Consideration | Why it matters |
|---|---|
| Pilot against ground truth before trusting it at scale | Run the tool's suggestions alongside human dispositions for a defined period and compare them directly, rather than assuming vendor-claimed accuracy transfers to your environment. |
| Keep a human decision point for closes and escalations | Anything that ends an alert's lifecycle, benign close or IR escalation, should require a human action to actually execute, not just a suggestion an analyst can ignore by doing nothing. |
| Log AI-suggested dispositions separately from human ones | Auditing the two streams independently is the only way to measure drift, bias, or a slow decline in suggestion quality over time, rather than discovering it after an incident. |
| Don't let copilots replace learning the query language | Analysts who never learn the underlying syntax cannot verify what a copilot ran or catch a misinterpreted request, a gap covered in more depth in chapter 2 of the Detection Engineering module. |
None of these considerations require slowing adoption to a crawl. They require building in the checkpoints that let a SOC find out it is wrong before that wrongness compounds into a missed intrusion or a burned-out team that has quietly stopped reading what the AI hands them.
Key Takeaways
- UEBA and unsupervised anomaly detection baseline normal behavior per entity and flag statistically significant deviations, catching novel behavior signatures miss, at the cost of a higher false-positive rate, especially early in baselining or in high-variance environments.
- AI-assisted triage tools summarize, enrich, and propose a disposition for incoming alerts, augmenting Tier 1 work rather than replacing it, because the cost of an incorrectly auto-closed true positive is too high to automate away entirely.
- SOC copilots translate natural-language questions into queries, lowering the barrier to routine lookups, but an analyst who cannot read the underlying query has no way to verify the translation was correct.
- Hallucination, missing organizational context, and automation bias are structural limitations, not temporary bugs, which is why human-in-the-loop review stays standard for consequential dispositions.
- Adopting these tools responsibly means piloting against ground truth, keeping a human checkpoint on closes and escalations, auditing AI dispositions separately, and not letting copilots substitute for learning the query language underneath them.
Knowledge Check
Click an answer to reveal the explanation.
Compared to a well-tuned signature-based detection, what is the core tradeoff of unsupervised anomaly detection like UEBA?
Why does human review of AI-suggested alert dispositions remain standard practice rather than being replaced by full automation?
An analyst has accepted a SOC copilot's "likely benign" recommendation correctly for months, and now reviews each new recommendation with less scrutiny than they used to, assuming the tool is probably right again. What is this pattern called?