CHAPTER 07 35 MIN READ ADVANCED

Detection-as-Code and the Rule Lifecycle

A detection rule that exists only inside a SIEM console is one bad click away from disappearing, and one undocumented edit away from becoming a mystery nobody on the team can explain. Detection-as-code borrows the discipline software engineering already worked out for exactly this problem: version control, peer review, automated testing, and a defined path from proposal to production. This chapter walks through what that discipline looks like when the "code" is detection logic rather than an application, and what happens at the other end of a detection's life when it's time to retire it.

detection-as-codeversion controlpeer reviewrule lifecycle
Building on earlier chapters: chapter 1 introduced detection-as-code in passing as a maturity signal for a detection program. This chapter is the deep dive that was promised there. It also assumes the testing practices from chapter 4 (backtesting, atomic testing) as a starting point: this chapter is about wrapping process, review, and automation around that testing rather than reintroducing it.

From Ad Hoc Rules to Detection-as-Code

Most detection programs start the same way. An analyst notices a gap, writes a query directly in the SIEM's rule editor, saves it, and moves on. For a handful of rules this works fine. There's no process overhead, the person who wrote the rule remembers why, and if something looks wrong they just open the console and fix it.

Why Ad Hoc Editing Breaks Down

The trouble is that this approach doesn't scale with the size of the rule set or the size of the team. Once a program has a few hundred detections maintained by a rotating group of engineers, the ad hoc model starts to fail in predictable ways:

  • Nobody outside the rule editor's change log (if the platform even keeps one) can tell who changed a piece of logic, when, or why.
  • A tuning change that quietly narrows a rule's coverage can sit unnoticed for months.
  • If an edit breaks something, there's often no clean way to see what the rule looked like before the change, let alone revert to it.
  • The reasoning behind a rule, why it exists, what it's meant to catch, why a particular exclusion was added, lives only in the head of whoever wrote it. When that person moves teams or leaves, the reasoning leaves with them.

What Detection-as-Code Means

Detection-as-code is the name for treating every detection rule the way a software team treats a piece of source code: stored in a repository, changed through a defined process, reviewed by a second person, tested before it ships, and reversible if it turns out to be wrong. None of the individual pieces are exotic. What changes is that they're applied consistently, as policy, rather than left to whichever engineer happens to be disciplined that week.

Note: Detection-as-code is a practice, not a specific product. It can be implemented with a SIEM's native rule format, with a portable format like Sigma, or with a mix of both, as long as the underlying logic lives in a version-controlled repository rather than only inside a vendor console.

Version Control for Detections

How Version Control Works Here

The foundation of detection-as-code is storing detection logic in a git repository instead of, or in addition to, a SIEM's built-in editor. This applies whether the logic is written natively in something like KQL or SPL, or expressed in a portable format like Sigma YAML that gets translated into a target platform's syntax at deployment time. The rule file itself might be short: a query, a title, a description, some metadata. What matters is that it lives as a tracked file alongside every other detection, with its full change history attached.

Three Benefits

Three benefits fall out of this almost immediately:

  • First, every change to a detection carries a commit: who made it, when, and, if commit messages are written with any care, why. A tuning change that added an exclusion six months ago is no longer a mystery; it's a line in `git log` with a message explaining the false-positive source it was written to address.
  • Second, changes can be developed and tested on a branch before they ever touch the production rule. An engineer can propose a new detection or a modification to an existing one without any risk to what's currently running, because the change doesn't take effect until it's merged.
  • Third, and often the most immediately useful in practice, rollback becomes trivial. If a deployed change turns out to be too noisy or, worse, silently breaks coverage, reverting to the last known-good version is a matter of checking out an earlier commit rather than trying to reconstruct what the rule used to say from memory or screenshots.

What Goes Into the Repository

A detection-as-code repository typically holds more than the raw query. Each rule's file, or an accompanying metadata block, usually carries the ATT&CK mapping from chapter 3, a short statement of the hypothesis the rule is built on, and enough context that someone unfamiliar with the rule can understand its intent without asking the original author. That context is what turns "a folder full of queries" into an actual maintainable system.

Peer Review Before Deployment

Version control on its own only gets a detection program part of the way there. The other piece is peer review: a second engineer looking at a proposed detection, or a proposed change to an existing one, before it merges into the branch that feeds production. This follows the same pull request or merge request pattern that's standard in software development, applied to detection logic instead of application code.

The Review Checklist

A good detection review isn't a rubber stamp. There's a specific, fairly consistent set of things a reviewer should be checking:

  • Does the logic actually match the stated hypothesis and the ATT&CK technique it claims to map to (chapter 3)? It's easy for a query to drift from its original intent during iteration, catching something adjacent to the hypothesis rather than the hypothesis itself.
  • Was the detection tested the way chapter 4 describes, with a backtest against historical data and, where applicable, an atomic test against a simulated technique? A reviewer should expect to see evidence of that testing, not just take the author's word for it.
  • Is the expected false-positive rate reasonable for the environment the rule will run in, and has the author accounted for the noisy legitimate activity that's likely to trigger it?
  • Is the rule documented well enough that someone other than the author could maintain it later: understand what it does, why it exists, and what to check first if it starts misbehaving?

Beyond the Checklist

The review step also does something less obvious but just as valuable: it spreads ownership. A detection that only one person has ever looked at is a detection only one person can confidently change. A detection that's been through review has at least two people who understand it, which matters enormously the first time that rule needs an urgent fix and its original author is unavailable.

Automated Testing Pipelines

Chapter 4 covered backtesting and atomic testing as things an engineer does by hand when validating a detection before deployment. In a detection-as-code workflow, those same tests can be wired into a CI/CD-style pipeline, similar in shape to the kind GitHub Actions or GitLab CI provide for application code, so they run automatically every time a change is proposed rather than only when someone remembers to run them.

The value of automating this isn't that it replaces human judgment, it's that it catches regressions consistently and immediately. A manual test that an engineer is supposed to run before every change is a test that will, sooner or later, get skipped under deadline pressure. An automated pipeline runs the same backtest and atomic test suite every single time, without needing anyone to remember, and flags a failure before a human reviewer even opens the pull request.

The Pipeline

1
Propose Change
Engineer opens a pull request with a new or modified detection.
→
2
Automated Test Run
Pipeline runs the backtest and/or atomic test suite against the change automatically.
→
3
Peer Review
A second engineer reviews logic, mapping, and documentation once tests pass.
→
4
Merge
Approved change is merged into the branch that represents production state.
→
5
Automated Deploy
Pipeline pushes the merged rule to the live detection platform.

Where Pipelines Stall

Where a pipeline stalls, the reason is usually infrastructure, not intent: getting a representative test data set wired into the pipeline, or building an atomic test harness that can run unattended. Those are worth the setup cost. A pipeline that only runs tests when someone remembers to trigger them manually is barely better than no pipeline at all.

Documentation Standards

A detection-as-code repository is only as maintainable as the documentation that travels with each rule. At minimum, every detection's file or accompanying metadata should carry:

FieldWhy it matters
HypothesisThe behavior the detection is built to catch, stated in plain language, so a future maintainer understands intent rather than just reading the logic.
ATT&CK mappingThe technique or sub-technique from chapter 3, keeping the rule tied into the coverage picture from chapter 6.
Known false-positive sourcesLegitimate activity that's expected to trigger the rule, and any tuning already applied to account for it, so the next person doesn't rediscover the same noise from scratch.
Change historyA record of what's changed since the rule was first written and why, beyond what a bare commit log conveys on its own.

None of this is bureaucracy for its own sake. It's what makes a detection maintainable by someone other than the person who wrote it. In a team with any amount of turnover, and most teams have some, that's the difference between a rule that stays useful for years and a rule that gets quietly disabled the first time it misbehaves because nobody left on the team understands what it was for.

Deprecation and Retirement as Part of the Lifecycle

A detection's lifecycle doesn't end when it stops being useful, it ends when it's deliberately removed. That distinction matters more than it sounds like it should.

Why Dead Rules Accumulate

Detections that quietly stop mattering, because the technology they targeted was retired, or the technique they covered was superseded by something else, tend to just sit in production indefinitely. Nobody actively decides to keep them; they're simply never actively removed.

Over time a rule set accumulates dead weight: detections that consume review and maintenance attention without contributing meaningful coverage, and in some cases detections whose false-positive load costs analyst time for no corresponding benefit.

Deliberate Removal

A mature detection-as-code practice treats retirement with the same rigor as deployment. Deprecating a rule goes through the same pull request and review pattern as adding one:

  • Someone proposes the removal and states the reasoning.
  • A reviewer confirms the rule really is redundant or obsolete before it comes out.
  • The removal itself gets documented in the same change history that tracked the rule's life, so if a question comes up later about why a given detection no longer exists, the answer is on record rather than left to memory.

Checking for Coverage Gaps

Just as importantly, retiring a detection should trigger a look back at the ATT&CK coverage mapping from chapter 6. If the rule being removed was the only coverage for a given technique, that removal opens a coverage gap, and someone needs to own the decision about whether that gap gets backfilled with a replacement detection or accepted as a known limitation.

Retirement handled this way keeps the rule set honest: what's running in production is what the team actually intends to be running, not an accumulation of history nobody has the context to clean up.

Key Takeaways

  • Ad hoc rule editing breaks down as a detection program grows: no history, no easy rollback, and knowledge trapped in one person's head.
  • Storing detection logic in a version-controlled repository gives every change a record of who/when/why, lets changes be branched and tested before they reach production, and makes rollback straightforward.
  • Peer review before deployment checks that a detection matches its hypothesis and ATT&CK mapping, was tested per chapter 4's process, has a reasonable false-positive risk, and is documented for future maintainers.
  • Automated testing pipelines run backtests and atomic tests on every proposed change, catching regressions before a human reviewer even looks at the pull request.
  • Documentation traveling with each rule (hypothesis, ATT&CK mapping, known false positives, change history) is what makes a detection maintainable by someone other than its original author.
  • Retirement deserves the same rigor as deployment: deliberate removal, documented reasoning, and a check against the ATT&CK coverage map for any gap the removal opens up.

Knowledge Check

Click an answer to reveal the explanation.

What is the main problem with editing detection rules directly in a SIEM console with no surrounding process, once a detection program grows past a handful of rules?

Ad hoc editing has no built-in history, review, or rollback path. As the rule count and team size grow, that gap turns into untraceable changes, silently broken detections, and knowledge that only exists in one person's memory.

In a detection-as-code peer review, which of the following is a reviewer specifically expected to check?

A meaningful review confirms the logic actually matches its stated hypothesis and ATT&CK mapping, that it was tested (backtest and/or atomic test), that its false-positive risk is reasonable, and that it's documented well enough for someone else to maintain later.

Why does retiring a detection deserve the same rigor as deploying one, rather than just disabling and forgetting it?

If a retired rule was the only coverage for a given technique, its removal creates a gap in the ATT&CK coverage mapping from chapter 6. Handling retirement deliberately, with documented reasoning and a coverage check, keeps that gap visible instead of letting it go unnoticed.