Testing and Validation
A detection rule that has never been tested is a hypothesis with extra steps. It might work exactly as intended, or it might sit silent forever, or it might fire on every login in the environment starting at 9 a.m. Monday. The only way to know which of those futures you're looking at is to test the logic before it reaches a real alert queue. This chapter covers how to validate a detection against data that already exists, how to generate the behavior it's meant to catch on demand, and how to read the results honestly enough to decide whether the rule is ready.
Why a Detection Needs Testing Before Deployment
Detection logic that reads correctly is not the same thing as detection logic that behaves correctly. A query can be syntactically valid, reference a field that genuinely exists in the schema, and still fail in production because the assumption behind it was wrong. On paper the rule is airtight. In practice it never fires.
- The log source doesn't populate that field under the conditions you expected.
- The process name is logged in a different case.
- The event only gets generated when a specific audit policy is enabled, and nobody enabled it.
The Opposite Failure: Too Broad to Be Useful
The opposite failure is just as common and arguably more damaging to a detection program's credibility: the condition is written too loosely, and it matches a huge slice of ordinary, legitimate activity. A rule looking for a scheduled task created by a non-standard account might be scoped correctly for a server fleet and catastrophically broad for a helpdesk team that provisions dozens of scheduled tasks a day as part of normal onboarding.
Deployed untested, that rule doesn't produce zero value, it produces negative value: it trains the analyst rotation to expect noise from this alert and to triage it without real attention, which is exactly the condition under which a genuine positive gets waved through.
Both failure modes are invisible from the query editor. The only way to surface either one is to run the logic against something real: data that already happened, or behavior generated specifically to test it. Testing isn't a formality tacked onto the end of the detection engineering workflow, it's the step that turns a plausible-looking rule into one you can actually stand behind.
Testing Against Historical Data
The first and cheapest test available to a detection engineer is a backtest: taking the draft detection logic and running it against a window of log data that already exists in the SIEM or data lake, rather than waiting for it to run live going forward.
Most platforms that support scheduled detection rules also let you run the same query manually against an arbitrary time range, which makes this step a matter of picking a reasonable lookback window (a few days to a few weeks, depending on data volume and how variable the target behavior is) and reviewing what comes back.
Backtesting answers a narrower question than atomic testing does, but it answers it fast: if this rule had been live for the last two weeks, how often would it have fired, and on what? A rule that returns zero results over two weeks of real traffic isn't necessarily broken, the behavior it targets may simply be rare, but a rule that returns several hundred results is telling you something concrete about scope.
- They're all coming from one specific service account doing something routine.
- They cluster around a particular business unit's normal workflow.
- There's a shared parent process that the detection never accounted for.
What backtesting is good at
Real historical data reveals false-positive sources that a detection engineer would rarely think to test for manually, because it reflects the actual mess of a production environment: legacy scripts, third-party tooling, IT automation, and edge cases baked in over years of operational history. No amount of imagining "what could cause false positives" from a desk replaces the signal of watching the logic run against what the environment actually produced.
What backtesting can't tell you
A backtest can only surface behavior that already occurred and was logged during the window you tested. If the technique the detection targets hasn't happened in your environment recently, or ever, a clean backtest tells you nothing about whether the logic actually fires when it should. A rule can pass a backtest with zero false positives and still be completely broken, silently failing to match the one thing it was built to catch. That gap is what the next section addresses.
Atomic Testing with Purpose-Built Simulation
Backtesting tells you how a detection behaves against activity that has already happened. It says nothing about whether the detection would catch the attacker behavior it was actually designed for, unless that behavior happened to occur, and get logged, during the window you tested. For most techniques, especially anything reasonably specific, that's not a safe assumption.
The fix is to stop waiting for the behavior to occur naturally and generate it on purpose, in a controlled way, and watch whether the detection fires.
What Atomic Red Team Does
This is what atomic testing means in a detection engineering context, and the best-known tool for it is Atomic Red Team, an open-source library maintained by Red Canary. It's built around small, discrete test scripts, each one mapped to a specific MITRE ATT&CK technique, designed to safely reproduce a narrow piece of attacker behavior (a particular command, a particular registry modification, a particular process relationship) without carrying out an actual attack.
Run one of these tests in a lab or sandboxed environment that feeds the same log sources your detection reads from, and you get a clean, known-cause event: you know exactly what happened, when it happened, and which technique it corresponds to.
From Guess to Direct Check
The value of that clean event is that it turns "does this detection fire on the right thing" from a guess into a direct check. Run the atomic test, then go look at whether the detection logic matched the resulting log data. If it fired, you've confirmed the rule works against the actual behavior it targets, not just against your mental model of that behavior. If it didn't fire, you've found the gap while the rule is still a draft, not after it's been live and silent for months.
Keep the test controlled
Atomic tests should run in an isolated lab, a dedicated test host, or another environment explicitly set aside for this purpose, never against arbitrary production systems. The goal is a clean signal you can attribute with certainty, and a segregated environment also keeps the exercise from generating noise or risk for people who have no idea a test is running.
Validating True Positive and False Positive Rates
Backtesting and atomic testing answer two different halves of the same question, and put together, they give a detection engineer a rough but grounded read on the rule's real-world behavior.
- Atomic test (true positive check): does the detection fire when the exact behavior it's meant to catch actually happens?
- Backtest (false positive estimate): how often does the same logic fire on activity that has nothing to do with that behavior?
Neither one alone is enough. A rule that fires reliably on the atomic test but also fires constantly in the backtest is technically working but operationally unusable. A rule that stays quiet in the backtest but never fires on the atomic test either isn't tuned, it's broken.
Treat this as a repeatable cycle rather than a one-time check, especially the first few times a piece of logic gets revised based on what testing surfaces.
"Decide" isn't a formality either. If the atomic test doesn't fire, the logic goes back to step one, there's no point tuning false positives on a rule that can't even catch its own target behavior. If the atomic test fires cleanly but the backtest volume is high, the rule is a candidate for the tuning process covered in the next chapter, not for immediate deployment.
Only a rule that clears both checks, confirmed true positive and a false positive rate the team is comfortable with, is genuinely ready to move toward production.
| Backtest result | Atomic test result | What it means |
|---|---|---|
| Low or no matches | Fires correctly | Strong candidate: rule appears scoped well and functionally correct |
| High volume of matches | Fires correctly | Logic works but is too broad: needs tuning before deployment |
| Low or no matches | Does not fire | Broken logic: likely a field mismatch or bad assumption about the data source |
| High volume of matches | Does not fire | Rule is both too broad and non-functional: back to the drafting stage |
Testing Safely, Outside Production
Every step above assumes the testing itself doesn't create risk for the people downstream of it. An untested rule should never be pointed at a live, unfiltered alert queue where a real analyst might see its output and act on it.
If the logic is wrong, an analyst could chase a false signal that wastes investigation time, or worse, could see a string of meaningless alerts from a rule they don't know is still in testing and start discounting alerts from that source by habit, which undermines the rule even after it's fixed.
Most SIEM and detection platforms offer some way to run logic live without exposing it to the production queue, and which one applies depends on what the platform supports:
- A staging or test queue. A separate alert destination, visible only to the detection engineering team, that mirrors production behavior without notifying the analyst rotation. New rules run here first and only get promoted to the production queue after they've cleared testing.
- Suppressed or silent deployment mode. Many platforms support deploying a rule in a mode that logs matches internally, sometimes called simulation mode or a dry-run, without generating a visible alert or notification to anyone. This lets the rule run against live, ongoing traffic (not just historical data) for a period of days or weeks so the engineer can see real-world match volume before flipping it to fully active.
- A dedicated testing index or workspace. Where the platform supports separate indices, workspaces, or rule sets, isolating test detections there entirely keeps them out of any path an analyst would normally look at, and out of dashboards or metrics that get reported on.
Whichever mechanism is available, the principle is the same: a detection stays in a space reserved for testing until it has cleared backtesting and atomic testing and someone has made a deliberate decision to promote it. That promotion step, and the tuning that often precedes it, is where the next chapter picks up.
Key Takeaways
- Detection logic fails in two directions: it never fires because of a bad assumption about the data, or it fires constantly because the condition is scoped too loosely. Testing exists to catch both before deployment.
- Backtesting runs draft logic against a window of historical data to surface obvious false-positive sources quickly, but it can only reveal behavior that already happened and was logged.
- Atomic testing, using a tool like Atomic Red Team, generates the target behavior on demand in a controlled environment, confirming a detection actually fires even for techniques the environment has never produced naturally.
- A rule isn't ready until both checks pass: a confirmed true positive from the atomic test and an acceptable false positive rate from the backtest.
- Untested logic never runs against a live, unfiltered production queue. Use a staging queue, a suppressed or silent deployment mode, or a dedicated testing index, depending on what the platform offers.
Knowledge Check
Click an answer to reveal the explanation.
A detection engineer backtests a new rule against two weeks of historical data and gets zero matches. What does this confirm on its own?
Why is Atomic Red Team useful for validating a detection, compared to relying on historical log data alone?
A new detection needs to run against live traffic for a few weeks before anyone trusts it with real alerts. Which approach fits that need without putting the rule in front of the analyst rotation?