CHAPTER 05 30 MIN READ INTERMEDIATE

False Positive Tuning

A detection rule that fires correctly in a test environment and a detection rule that survives contact with a live, messy production environment are not the same thing. The gap between them is false positives: alerts the logic technically matched but that don't represent the malicious activity the rule was written to catch. Tuning is the ongoing work of closing that gap without closing off the rule's ability to catch what it was actually built for. Done well, it turns a noisy, ignored detection into one analysts trust. Done poorly, or not at all, it does the opposite.

false positive tuningallowlistingalert fatiguetuning cadence
Where this picks up: This chapter assumes the detection has already passed the testing and validation stage from Chapter 4 and is now live in production, generating real alerts against real traffic. Tuning here is a response to real-world feedback, not a pre-deployment checklist item.

Why Detections Generate False Positives

A false positive isn't a sign the detection is broken in the sense of not working. It usually means the logic is working exactly as written, and the gap is between what was written and what the engineer actually intended to catch. Most false positives trace back to one of four root causes.

Root causeWhat's happening
Overly broad logicThe rule's conditions match a wider set of activity than the engineer had in mind. A pattern written to catch "process X launched with these arguments" ends up matching several legitimate call paths that happen to share the same argument structure.
Dual-use legitimate toolsThe technique being detected is also how administrators, backup agents, or monitoring software legitimately do their job. This is especially common in Windows and living-off-the-land style detections, where the same binary an attacker abuses for lateral movement is also what the help desk uses every day.
Unaccounted-for normal behaviorThe environment has a routine, expected process the engineer simply didn't know about when writing the rule, such as a scheduled maintenance script that happens to create a registry run key in a way that matches a persistence detection.
Data quality issuesA field doesn't mean what the engineer assumed it meant, or contains values the engineer didn't anticipate. This is the same field-selection risk covered in Chapter 3: build logic on a field you haven't verified, and the rule inherits whatever assumption was wrong.

These causes aren't mutually exclusive. A single noisy detection can be a bit of all four: logic that's slightly broader than needed, applied against a log source with a quirky field, in an environment that happens to run a dual-use admin tool nobody flagged during the requirements stage.

Diagnosing which one (or which combination) is driving the noise is the first step in tuning, because each root cause points toward a different fix.

Note: Before tuning anything, confirm the alerts really are false positives and not a true positive that looks unfamiliar. Pulling a sample of recent alerts and manually walking through what generated each one is worth the time before touching the logic.

The Cost of Not Tuning

An untuned detection has a cost that goes beyond the minutes an analyst spends closing each alert. If you've been through the triage material in the SOC Operations module, the term alert fatigue will already be familiar: the state where a high enough volume of low-value alerts erodes an analyst's attention and default response to a given alert type.

How Repetition Erodes Analyst Response

A detection that fires dozens of times a day for the same known-benign reason doesn't just waste time on those dozens of closures. It trains the analysts handling it. After enough repetitions of "this alert, closed as benign," the human response becomes automatic: see the alert, recognize the pattern, close it, move on.

That's a rational adaptation to a noisy queue, and it's also exactly the condition under which a real true positive slips through. The one time that specific alert represents genuine malicious activity, it looks identical to the hundred times before it that didn't, and it gets the same reflexive close.

This is why tuning isn't a nice-to-have polish step. An untuned high-fidelity detection can end up worse than no detection at all, because it consumes analyst trust and attention while providing a false sense that the technique is covered.

Tuning Approaches

Once the root cause is understood, the fix usually comes down to one of four levers. They aren't interchangeable: picking the wrong one either fails to fix the noise or quietly opens a coverage gap.

LeverWhat it doesWhen it's the right tool
Allowlisting / exceptionsExcludes a specific, identified source: a host, user account, process path, or file hash known to be legitimate.The false positive source is a small, stable, well-identified list (a backup server, a specific service account) rather than a broad category of activity.
Threshold adjustmentRaises the frequency or count requirement so normal-volume activity falls below the alerting bar.The false positives come from legitimate activity that happens at a lower volume than the malicious activity the rule targets, and volume is a meaningful signal for that technique.
Logic narrowingAdds an additional condition to the existing pattern that separates legitimate instances from malicious ones sharing the same base signature.Legitimate and malicious activity both match the current logic, but differ on some other observable attribute (a parent process, a command-line flag, a destination) that the rule doesn't currently check.
Time-window adjustmentChanges how far apart correlated events can occur and still be treated as related for a multi-event detection.The correlation window is too wide or too narrow for how the underlying activity, legitimate or malicious, actually unfolds in this environment.

Choosing between them

Trade-offs:
  • Allowlisting is the most surgical option and the easiest to reason about, but it doesn't scale if the list of exceptions keeps growing: a rule with fifty exceptions is a sign the underlying logic needs narrowing, not more exceptions.
  • Threshold and time-window adjustments are attractive because they're a single number to change, but they can quietly raise the bar for detecting the real technique too, so they need to be checked against the lower bound of what real malicious activity would actually produce.
  • Logic narrowing takes more work up front but tends to produce the most durable fix, because it's addressing the actual distinguishing signal rather than a proxy for it.

Tuning vs. Suppression vs. Retirement

Not every noisy detection deserves the same response. Some should be tuned, some should keep running quietly in the background without alerting, and some should be turned off entirely. Working through the same few questions in order keeps that decision consistent across a detection engineering team instead of depending on who happens to be on shift.

1
Is the false positive pattern narrow and identifiable?
If the noise traces to a specific, describable source and excluding it doesn't remove real coverage: tune it, using whichever lever from the table above fits the cause.
→
2
Is the detection still useful for investigation, just too noisy to alert on?
If the logic reliably captures relevant activity but can't be tuned down to an alertable rate without losing that value: suppress it. Keep it logging and available as investigative context, but stop pushing it to the alert queue.
→
3
Can the false positive rate come down without gutting real coverage?
If every tuning attempt that reduces noise also removes the detection's ability to catch the real technique, or the underlying technique or environment has changed enough that the rule no longer applies: retire it. A detection that can't be tuned into something trustworthy is worse than no detection.

The order matters. Jumping straight to retirement because a detection is annoying skips the chance to recover a rule that just needs a narrower condition. Jumping straight to suppression to make the noise go away, without first attempting to tune, can hide a rule that would have been perfectly usable with one added condition.

Tuning Cadence and Ownership

Tuning isn't something that happens once, right after a detection goes live, and then is done. Environments keep changing after a rule ships: new software gets rolled out, a team adopts a new admin tool, an infrastructure migration changes what normal traffic looks like. None of that requires the detection's logic to be wrong on day one; it just means the detection drifts out of tune as the environment around it moves.

Two things keep that drift from accumulating unnoticed.

Regular review cadence

The first is a regular review cadence: a scheduled point, not just a reactive one, where recently deployed and previously tuned detections get reassessed against current alert volume and false positive rate.

Clear ownership

The second is clear ownership. Tuning feedback that has no assigned owner tends to sit in a queue or a Slack thread indefinitely, waiting for "whoever notices" to act on it. Assigning a specific person or rotation responsibility for acting on tuning feedback, separate from whoever originally wrote the rule, keeps that feedback loop from stalling out as the team and the rule set both grow.

Key Takeaways

  • False positives usually trace to one of four root causes: overly broad logic, dual-use legitimate tools, unaccounted-for normal behavior, or data quality issues.
  • An untuned detection costs more than analyst time: it conditions analysts to dismiss that alert type reflexively, which is how a real true positive gets missed.
  • The main tuning levers are allowlisting, threshold adjustment, logic narrowing, and time-window adjustment, each suited to a different type of false positive pattern.
  • A noisy detection should be tuned if the false positive pattern is narrow, suppressed if it's still useful for context but too noisy to alert on, or retired if its false positive rate can't come down without losing real coverage.
  • Tuning is ongoing, not a one-time step. Environments drift, so detections need a regular review cadence and a clearly assigned owner to act on tuning feedback.

Knowledge Check

Click an answer to reveal the explanation.

A detection for a persistence technique keeps firing on a scheduled maintenance script that legitimately writes to the same registry location the rule watches. Which root cause best describes this?

The rule's logic and the data are both fine here: it's correctly matching a registry write in the expected location. The gap is that the engineer didn't know about this specific scheduled script when writing the rule, which is the definition of unaccounted-for normal behavior rather than a logic or data problem.

A detection has grown to over sixty individual host and process exceptions over time, and the exception list keeps growing every month. What does that pattern usually indicate?

A short, stable exception list is a normal and reasonable use of allowlisting. A list that keeps growing is a sign the rule is matching a broad category of legitimate activity rather than a handful of known exceptions, and the more durable fix is usually to narrow the logic so it stops matching that category in the first place.

A detection reliably captures relevant activity and is valuable during investigations, but its false positive rate is too high to alert on directly, and no combination of tuning levers brings that rate down without losing real coverage. What's the appropriate response?

This is the definition of a suppression case: the detection is still useful, just not at alert-worthy fidelity. Retiring it would throw away a detection that has real investigative value, and continuing to alert on it feeds the alert fatigue problem covered earlier in the chapter.