CHAPTER 08 30 MIN READ ADVANCED

Metrics and Detection Engineering Maturity

Chapter 1 opened this module with a lifecycle: hypothesis, logic, test, tune, deploy, maintain. Seven chapters later, that lifecycle has been through every stage in detail, from writing the first Sigma rule to keeping a rule catalog from rotting. This closing chapter asks a different kind of question: how does a program know whether any of it is actually working? Detection engineering has its own set of metrics, separate from the alert-handling metrics a SOC already tracks, and reading them correctly is what separates a program that improves over time from one that just accumulates rules. This chapter covers those metrics, draws a clear line between detection engineering measurement and SOC operations measurement, and closes the module by tying the eight chapters into one arc.

detection metricsfalse positive rateprogram maturitycoverage percentage
Capstone chapter: this is the last chapter in the module, and it deliberately loops back to chapter 1. Everything here is about measuring whether the lifecycle you learned in chapter 1, and practiced in chapters 3 through 7, is holding up at program scale rather than just working for one rule at a time.

Detection Engineering's Own Metrics

A SOC produces plenty of numbers, but most of them describe what happens after a detection fires: how long until someone looked at the alert, how long until it was resolved, how many alerts a shift handled. Those numbers say almost nothing about whether the detections themselves are any good.

Detection engineering needs its own metrics, ones that describe the quality of the detection logic and the health of the catalog it lives in, not the response process built on top of it.

False positive rate per detection

Chapter 5 covered the tuning workflow in depth: track dispositions, feed them back into the rule, adjust thresholds or add exclusions, and repeat. The false positive rate for a given detection is the running tally of that work.

Definition: False positive rate per detection is benign-classified alerts divided by total alerts for that rule, over a defined window.

A single rule with a high false positive rate is a tuning problem. A catalog where most rules sit at a high false positive rate is a process problem: either tuning isn't happening on a cadence, or new rules are shipping without enough testing before they go live.

This number is only useful per rule, or grouped by a meaningful category (technique, data source, rule author). A single blended false positive rate across the whole catalog hides exactly the rules that need attention behind the ones that don't.

Time from hypothesis to deployed detection

Chapter 1 laid out the lifecycle stages; this metric measures how efficiently a team actually moves through them. It's the elapsed time between a hypothesis being written down (often coming out of a threat hunt, as the closing section below revisits) and a tested, tuned detection reaching production.

A program where this interval keeps growing usually has a bottleneck somewhere in build, test, or review, and it's worth knowing which stage before assuming the whole lifecycle is broken. A program where the interval is suspiciously short is worth checking for the opposite problem: stages getting skipped rather than accelerated.

ATT&CK coverage percentage, with depth

Chapter 6 introduced ATT&CK-mapped coverage and warned against treating the raw percentage as a scoreboard. That warning carries forward here as a measurement, not just a framing point: a coverage percentage on its own answers "how many techniques have at least one detection," and nothing about whether that detection is any good, how many known variations of the technique it actually catches, or whether it has ever fired on a real technique execution during testing.

Better metric: Coverage percentage paired with the none/partial/broad depth rating chapter 6 introduced is a far more honest number than the percentage alone.

Detection staleness

Chapter 5's tuning cadence point extends into a metric here: how many deployed detections have gone past their expected review interval without being touched. A rule that hasn't been reviewed, tuned, or re-tested in a long stretch isn't necessarily broken, but nobody currently knows whether it is.

Staleness is a leading indicator: a catalog with a growing stale-rule count is a catalog that will eventually surprise someone, either with a rule that silently stopped firing after an environment change, or one that's been quietly generating noise nobody's looked at.

Note: None of these four metrics is meaningful in isolation for very long. A team can chase false positive rate down by narrowing a rule until it barely fires, which looks great on that metric and terrible for coverage depth. Read them together.

How This Connects to SOC-Wide Metrics

Readers who came through the SOC Operations module will recognize metrics like mean time to detect and mean time to respond. Those are real, useful numbers, and detection engineering work directly influences them. But they belong to SOC operations, not to detection engineering, and mixing the two up is a common source of confused reporting.

Mean time to detect, as an example

Take mean time to detect. It measures the gap between when malicious activity started and when the SOC first noticed it. Whether a detection existed for that technique, and whether it fired correctly when the activity occurred, is one of the biggest levers on that number. If detection engineering has a coverage gap on a technique, or a rule for that technique is disabled because it was too noisy, mean time to detect for that scenario will suffer.

But the metric itself is measuring the organization's overall detection-to-response pipeline, including alert triage, escalation, and analyst workload, none of which detection engineering controls.

Where the boundary sits

The cleanest way to hold the boundary: detection engineering metrics describe the quality and coverage of the detections themselves, independent of what happens after they fire. SOC operations metrics describe what happens once an alert reaches a human, from triage through resolution.

A detection engineering team can point to a strong false positive rate, solid coverage depth, and a low staleness count and still see MTTD stay flat, because the bottleneck moved into triage capacity or escalation process. Conversely, a SOC can tighten its response process and improve MTTR without any change to the underlying detection catalog.

Watch for: Reporting that conflates the two, crediting a detection engineering effort for a response-time improvement it didn't cause, or blaming detection engineering for a triage backlog, erodes trust in both sets of numbers.
MetricOwned byAnswers
False positive rate per detectionDetection engineeringIs this specific rule tuned well?
ATT&CK coverage + depthDetection engineeringWhat can we see, and how well?
Detection stalenessDetection engineeringIs the catalog being maintained?
Mean time to detectSOC operationsHow fast did we notice, end to end?
Mean time to respondSOC operationsHow fast did we act once we noticed?

What Separates a Mature Detection Program From an Ad Hoc One

There's no single external scorecard that reliably tells a team where it stands on detection engineering maturity, and treating a borrowed number from an unrelated discipline as if it were one usually does more harm than good.

What's more useful, and more honest, is thinking about maturity as a handful of dimensions a program can assess for itself. None of these require a formal audit; each one is answerable by looking at how the team actually works day to day.

Consistency

Does every detection, regardless of who wrote it, go through the same lifecycle: hypothesis, logic, test, tune, deploy, and maintain, with peer review applied at every step? In an ad hoc program, quality tracks individual engineers. The senior person's rules are solid and the newer hire's rules are inconsistent, not because of a skill gap that can't be closed, but because there's no shared process forcing both through the same bar.

Documentation

Chapter 7 covered this at the individual-rule level: can someone other than the original author read the documentation and understand what the rule does, why it exists, and how to maintain it. At the program level, this dimension asks whether that standard is actually enforced across the whole catalog, or whether it's aspirational for anything written more than a few months ago.

Measurement

Does the program track its own metrics from the first section of this chapter, or does it operate on the assumption that things are probably fine because nobody's complained? A team that can't currently answer "what's our false positive rate on the top ten noisiest rules" doesn't have a measurement practice yet, whatever else is going well.

Feedback loops

Tuning feedback from chapter 5 and coverage-gap analysis from chapter 6 are only useful if they turn into scheduled work. A mature program has a mechanism, a backlog, a recurring review meeting, whatever fits the team, that reliably converts "this rule is noisy" or "we have no coverage here" into an assigned task with a timeline.

An ad hoc program generates the same insights and then loses them; the gap gets noted once, in a chat message or a postmortem, and nobody revisits it.

DimensionAd hocMature
ConsistencyQuality depends on which engineer wrote the rule; no shared checklistEvery detection goes through the same build/test/review/deploy steps
DocumentationRules are self-explanatory only to their author, if thatAny team member can maintain a rule they didn't write
MeasurementHealth is assumed; no one can produce current false positive or staleness numbersMetrics are tracked routinely and reviewed, not just collected
Feedback loopsTuning notes and gap findings are mentioned once and forgottenFindings become tracked, assigned work with follow-through
Note: A program can be strong on one dimension and weak on another. It's common to see solid documentation habits paired with no real measurement practice, or good metrics paired with a feedback loop that quietly stalls. Assess the four separately rather than averaging them into one impression.

Common Measurement Pitfalls

Metrics can mislead as easily as they can inform, and detection engineering has a few recurring traps worth naming directly.

Watch for:
  • Coverage percentage as a vanity metric. Chapter 6 raised this first, and it's worth restating here because this is the chapter where numbers get reported upward: a high percentage of ATT&CK techniques with "coverage" says nothing about depth, and a program that optimizes for the percentage alone will end up with a lot of shallow, log-only entries rather than tested, reliable detections.
  • Raw detection volume as a proxy for program health. A catalog with more rules isn't automatically a stronger catalog. If a meaningful share of those rules are noisy, redundant with each other, or covering the same narrow slice of a technique from slightly different angles, the count is inflated without the coverage or quality to back it up.
  • Rules shipped per engineer per period as a performance measure. This is the most damaging pitfall because it works against everything the rest of this module built. Measuring engineers by output volume rewards writing more rules faster, which pulls directly against the testing rigor from chapter 4, the tuning discipline from chapter 5, and the documentation standard from chapter 7. A program that measures this way will get a faster-growing catalog and a worse one at the same time.

The common thread across all three: each one substitutes something easy to count for something harder to assess. Depth, quality, and rigor take more effort to measure than a percentage or a headcount, but they're the numbers that actually describe whether the program is working.

Closing the Module: The Full Arc

The module's arc, chapter by chapter

ChapterRole in the arc
1Introduced the detection engineering lifecycle as a whole
2Built the language literacy needed to read and write across query languages rather than being locked to one
3 and 4Narrowed to a single detection: building it, then proving it actually works before it ever reaches production
5Tuning a live detection over time
6Thinking about coverage across the whole catalog rather than one rule at a time
7Treating detection logic as code with a real lifecycle and real documentation
8 (this chapter)How a program measures whether all of that is actually paying off

That arc, foundations and language, then a single detection built and proven, then sustaining and scaling that work across a program, is the same shape most detection engineering practices grow into, whether or not a team ever names the stages explicitly.

Where this module fits

This module doesn't stand alone.

Related modules:
  • SOC Operations covers the platform and process layer this work eventually feeds into: how alerts get triaged, escalated, and resolved once a detection fires.
  • Threat Hunting covers where many of the hypotheses that start chapter 1's lifecycle actually come from, an analyst noticing something during a hunt that doesn't have a detection behind it yet.

Detection engineering sits between the two: it takes a hunting insight and turns it into something repeatable, and it hands that repeatable detection off to a SOC operations process built to act on it.

Key Takeaways

  • Detection engineering has its own KPIs: false positive rate per detection, time from hypothesis to deployed detection, ATT&CK coverage percentage paired with a depth rating, and detection staleness.
  • These metrics are distinct from SOC-wide metrics like MTTD and MTTR. Detection engineering measures the quality and coverage of detections; SOC operations measures what happens after an alert fires.
  • Program maturity is best assessed across four dimensions: consistency, documentation, measurement, and feedback loops, rather than reduced to a single borrowed score.
  • Coverage percentage alone, raw detection volume, and rules-shipped-per-engineer are all misleading measures of program health on their own.
  • The module's arc runs from foundations and language literacy, through building and proving a single detection, to sustaining and measuring that work at program scale, and it connects directly to the SOC Operations and Threat Hunting modules.

Knowledge Check

Click an answer to reveal the explanation.

A detection catalog shows a rising number of rules that haven't been reviewed or tuned past their expected interval. Which metric captures this directly?

Detection staleness tracks how many deployed detections have gone past their review or tuning interval. It's a leading indicator: the rules aren't necessarily broken, but nobody currently knows their state, which is exactly the risk chapter 5's tuning cadence was meant to prevent.

Mean time to detect (MTTD) improves after a detection engineering team closes a coverage gap. Which statement best describes the relationship?

Whether a detection existed and fired correctly is a major lever on MTTD, but MTTD measures the organization's full detection-to-response pipeline, including triage and escalation. It stays a SOC operations metric even when detection engineering work is a big part of why it moved.

A manager starts evaluating engineers by counting the number of detection rules each one ships per quarter. What is the main risk of this measurement approach?

Rules-shipped-per-period rewards speed and quantity. That pulls directly against the testing rigor from chapter 4, the tuning discipline from chapter 5, and the documentation standard from chapter 7, producing a catalog that grows faster while getting less reliable.