Metrics and Detection Engineering Maturity
Chapter 1 opened this module with a lifecycle: hypothesis, logic, test, tune, deploy, maintain. Seven chapters later, that lifecycle has been through every stage in detail, from writing the first Sigma rule to keeping a rule catalog from rotting. This closing chapter asks a different kind of question: how does a program know whether any of it is actually working? Detection engineering has its own set of metrics, separate from the alert-handling metrics a SOC already tracks, and reading them correctly is what separates a program that improves over time from one that just accumulates rules. This chapter covers those metrics, draws a clear line between detection engineering measurement and SOC operations measurement, and closes the module by tying the eight chapters into one arc.
Detection Engineering's Own Metrics
A SOC produces plenty of numbers, but most of them describe what happens after a detection fires: how long until someone looked at the alert, how long until it was resolved, how many alerts a shift handled. Those numbers say almost nothing about whether the detections themselves are any good.
Detection engineering needs its own metrics, ones that describe the quality of the detection logic and the health of the catalog it lives in, not the response process built on top of it.
False positive rate per detection
Chapter 5 covered the tuning workflow in depth: track dispositions, feed them back into the rule, adjust thresholds or add exclusions, and repeat. The false positive rate for a given detection is the running tally of that work.
A single rule with a high false positive rate is a tuning problem. A catalog where most rules sit at a high false positive rate is a process problem: either tuning isn't happening on a cadence, or new rules are shipping without enough testing before they go live.
This number is only useful per rule, or grouped by a meaningful category (technique, data source, rule author). A single blended false positive rate across the whole catalog hides exactly the rules that need attention behind the ones that don't.
Time from hypothesis to deployed detection
Chapter 1 laid out the lifecycle stages; this metric measures how efficiently a team actually moves through them. It's the elapsed time between a hypothesis being written down (often coming out of a threat hunt, as the closing section below revisits) and a tested, tuned detection reaching production.
A program where this interval keeps growing usually has a bottleneck somewhere in build, test, or review, and it's worth knowing which stage before assuming the whole lifecycle is broken. A program where the interval is suspiciously short is worth checking for the opposite problem: stages getting skipped rather than accelerated.
ATT&CK coverage percentage, with depth
Chapter 6 introduced ATT&CK-mapped coverage and warned against treating the raw percentage as a scoreboard. That warning carries forward here as a measurement, not just a framing point: a coverage percentage on its own answers "how many techniques have at least one detection," and nothing about whether that detection is any good, how many known variations of the technique it actually catches, or whether it has ever fired on a real technique execution during testing.
Detection staleness
Chapter 5's tuning cadence point extends into a metric here: how many deployed detections have gone past their expected review interval without being touched. A rule that hasn't been reviewed, tuned, or re-tested in a long stretch isn't necessarily broken, but nobody currently knows whether it is.
Staleness is a leading indicator: a catalog with a growing stale-rule count is a catalog that will eventually surprise someone, either with a rule that silently stopped firing after an environment change, or one that's been quietly generating noise nobody's looked at.
How This Connects to SOC-Wide Metrics
Readers who came through the SOC Operations module will recognize metrics like mean time to detect and mean time to respond. Those are real, useful numbers, and detection engineering work directly influences them. But they belong to SOC operations, not to detection engineering, and mixing the two up is a common source of confused reporting.
Mean time to detect, as an example
Take mean time to detect. It measures the gap between when malicious activity started and when the SOC first noticed it. Whether a detection existed for that technique, and whether it fired correctly when the activity occurred, is one of the biggest levers on that number. If detection engineering has a coverage gap on a technique, or a rule for that technique is disabled because it was too noisy, mean time to detect for that scenario will suffer.
But the metric itself is measuring the organization's overall detection-to-response pipeline, including alert triage, escalation, and analyst workload, none of which detection engineering controls.
Where the boundary sits
The cleanest way to hold the boundary: detection engineering metrics describe the quality and coverage of the detections themselves, independent of what happens after they fire. SOC operations metrics describe what happens once an alert reaches a human, from triage through resolution.
A detection engineering team can point to a strong false positive rate, solid coverage depth, and a low staleness count and still see MTTD stay flat, because the bottleneck moved into triage capacity or escalation process. Conversely, a SOC can tighten its response process and improve MTTR without any change to the underlying detection catalog.
| Metric | Owned by | Answers |
|---|---|---|
| False positive rate per detection | Detection engineering | Is this specific rule tuned well? |
| ATT&CK coverage + depth | Detection engineering | What can we see, and how well? |
| Detection staleness | Detection engineering | Is the catalog being maintained? |
| Mean time to detect | SOC operations | How fast did we notice, end to end? |
| Mean time to respond | SOC operations | How fast did we act once we noticed? |
What Separates a Mature Detection Program From an Ad Hoc One
There's no single external scorecard that reliably tells a team where it stands on detection engineering maturity, and treating a borrowed number from an unrelated discipline as if it were one usually does more harm than good.
What's more useful, and more honest, is thinking about maturity as a handful of dimensions a program can assess for itself. None of these require a formal audit; each one is answerable by looking at how the team actually works day to day.
Consistency
Does every detection, regardless of who wrote it, go through the same lifecycle: hypothesis, logic, test, tune, deploy, and maintain, with peer review applied at every step? In an ad hoc program, quality tracks individual engineers. The senior person's rules are solid and the newer hire's rules are inconsistent, not because of a skill gap that can't be closed, but because there's no shared process forcing both through the same bar.
Documentation
Chapter 7 covered this at the individual-rule level: can someone other than the original author read the documentation and understand what the rule does, why it exists, and how to maintain it. At the program level, this dimension asks whether that standard is actually enforced across the whole catalog, or whether it's aspirational for anything written more than a few months ago.
Measurement
Does the program track its own metrics from the first section of this chapter, or does it operate on the assumption that things are probably fine because nobody's complained? A team that can't currently answer "what's our false positive rate on the top ten noisiest rules" doesn't have a measurement practice yet, whatever else is going well.
Feedback loops
Tuning feedback from chapter 5 and coverage-gap analysis from chapter 6 are only useful if they turn into scheduled work. A mature program has a mechanism, a backlog, a recurring review meeting, whatever fits the team, that reliably converts "this rule is noisy" or "we have no coverage here" into an assigned task with a timeline.
An ad hoc program generates the same insights and then loses them; the gap gets noted once, in a chat message or a postmortem, and nobody revisits it.
| Dimension | Ad hoc | Mature |
|---|---|---|
| Consistency | Quality depends on which engineer wrote the rule; no shared checklist | Every detection goes through the same build/test/review/deploy steps |
| Documentation | Rules are self-explanatory only to their author, if that | Any team member can maintain a rule they didn't write |
| Measurement | Health is assumed; no one can produce current false positive or staleness numbers | Metrics are tracked routinely and reviewed, not just collected |
| Feedback loops | Tuning notes and gap findings are mentioned once and forgotten | Findings become tracked, assigned work with follow-through |
Common Measurement Pitfalls
Metrics can mislead as easily as they can inform, and detection engineering has a few recurring traps worth naming directly.
- Coverage percentage as a vanity metric. Chapter 6 raised this first, and it's worth restating here because this is the chapter where numbers get reported upward: a high percentage of ATT&CK techniques with "coverage" says nothing about depth, and a program that optimizes for the percentage alone will end up with a lot of shallow, log-only entries rather than tested, reliable detections.
- Raw detection volume as a proxy for program health. A catalog with more rules isn't automatically a stronger catalog. If a meaningful share of those rules are noisy, redundant with each other, or covering the same narrow slice of a technique from slightly different angles, the count is inflated without the coverage or quality to back it up.
- Rules shipped per engineer per period as a performance measure. This is the most damaging pitfall because it works against everything the rest of this module built. Measuring engineers by output volume rewards writing more rules faster, which pulls directly against the testing rigor from chapter 4, the tuning discipline from chapter 5, and the documentation standard from chapter 7. A program that measures this way will get a faster-growing catalog and a worse one at the same time.
The common thread across all three: each one substitutes something easy to count for something harder to assess. Depth, quality, and rigor take more effort to measure than a percentage or a headcount, but they're the numbers that actually describe whether the program is working.
Closing the Module: The Full Arc
The module's arc, chapter by chapter
| Chapter | Role in the arc |
|---|---|
| 1 | Introduced the detection engineering lifecycle as a whole |
| 2 | Built the language literacy needed to read and write across query languages rather than being locked to one |
| 3 and 4 | Narrowed to a single detection: building it, then proving it actually works before it ever reaches production |
| 5 | Tuning a live detection over time |
| 6 | Thinking about coverage across the whole catalog rather than one rule at a time |
| 7 | Treating detection logic as code with a real lifecycle and real documentation |
| 8 (this chapter) | How a program measures whether all of that is actually paying off |
That arc, foundations and language, then a single detection built and proven, then sustaining and scaling that work across a program, is the same shape most detection engineering practices grow into, whether or not a team ever names the stages explicitly.
Where this module fits
This module doesn't stand alone.
- SOC Operations covers the platform and process layer this work eventually feeds into: how alerts get triaged, escalated, and resolved once a detection fires.
- Threat Hunting covers where many of the hypotheses that start chapter 1's lifecycle actually come from, an analyst noticing something during a hunt that doesn't have a detection behind it yet.
Detection engineering sits between the two: it takes a hunting insight and turns it into something repeatable, and it hands that repeatable detection off to a SOC operations process built to act on it.
Key Takeaways
- Detection engineering has its own KPIs: false positive rate per detection, time from hypothesis to deployed detection, ATT&CK coverage percentage paired with a depth rating, and detection staleness.
- These metrics are distinct from SOC-wide metrics like MTTD and MTTR. Detection engineering measures the quality and coverage of detections; SOC operations measures what happens after an alert fires.
- Program maturity is best assessed across four dimensions: consistency, documentation, measurement, and feedback loops, rather than reduced to a single borrowed score.
- Coverage percentage alone, raw detection volume, and rules-shipped-per-engineer are all misleading measures of program health on their own.
- The module's arc runs from foundations and language literacy, through building and proving a single detection, to sustaining and measuring that work at program scale, and it connects directly to the SOC Operations and Threat Hunting modules.
Knowledge Check
Click an answer to reveal the explanation.
A detection catalog shows a rising number of rules that haven't been reviewed or tuned past their expected interval. Which metric captures this directly?
Mean time to detect (MTTD) improves after a detection engineering team closes a coverage gap. Which statement best describes the relationship?
A manager starts evaluating engineers by counting the number of detection rules each one ships per quarter. What is the main risk of this measurement approach?