Detection Engineering Foundations
Detection engineering is the discipline of turning a belief about how an attacker might behave into a working, tuned, production alert. It sits next to the SOC and next to threat hunting, but it is not the same job: analysts triage what fires, hunters go looking for what nobody built a rule for yet, and detection engineers build the rule so the next occurrence gets caught automatically. This chapter defines the role, walks through the lifecycle a detection moves through from idea to maintenance, and introduces the detection-as-code mindset that keeps a growing rule set from collapsing under its own weight. Everything here is the overview: chapters 2 through 7 each take one stage of this lifecycle and go deep.
What a Detection Engineer Actually Does
Strip away the job title and the work is narrow and concrete: take a hypothesis about attacker behavior and turn it into something a SIEM or EDR can evaluate against live data, reliably, without burying analysts in noise. That hypothesis rarely comes from nowhere.
- A threat intel report describing a new technique used by an actor relevant to the organization.
- A threat hunter who found something suspicious once and now wants it caught automatically.
- A gap identified while mapping coverage against MITRE ATT&CK.
- A postmortem where the honest answer to "why didn't we catch this" was "nobody wrote a rule for it."
From there the detection engineer's job is to translate that idea into logic: pick the right data source, write the query or rule in whatever language the platform uses (KQL, SPL, Sigma, or a vendor-specific syntax), test it against both malicious and benign activity, tune out the false positives without tuning out the true positives, document what the rule does and why, and ship it.
Then the job continues, because a detection that works today can go stale as the environment changes, as the attack technique evolves, or as a new benign process starts triggering it.
How This Differs From Adjacent Roles
Detection engineering overlaps with SOC analyst work and threat hunting in the skills required, mostly log analysis, query writing, and attacker behavior knowledge, but the three roles produce different outputs and operate on different timelines.
| Role | Primary Question | Primary Output | Timing |
|---|---|---|---|
| Detection Engineer | How do we catch this reliably, every time, going forward? | A tuned, production detection rule | Ongoing, proactive, builds lasting capability |
| SOC Analyst (Tier 1-3) | Is this specific alert a real problem, and what do we do about it? | A triaged and resolved alert or incident | Reactive, works the queue as it arrives |
| Threat Hunter | Is this activity already happening in our environment undetected? | A hunt finding, which may become a new detection request | Periodic, hypothesis-driven, proactive but not continuous |
In practice these roles feed each other constantly. A hunter's finding becomes a detection engineer's backlog item. An analyst's repeated manual triage of the same alert pattern becomes a tuning request. A detection engineer's new rule becomes the next alert an analyst has to learn to triage.
The SOC Operations module covers the analyst tiers in depth, and the Threat Hunting module covers hypothesis-driven hunting; this module focuses on the engineering side of that loop.
The Detection Engineering Lifecycle
A detection does not go from idea to production in one step, and it does not stop being work once it is deployed. Most teams that do this well recognize six stages, whether or not they call them by these names. This chapter introduces all six at a glance; chapters 3 through 7 each dedicate space to one or more of them.
Note that the lifecycle loops back on itself. A rule in the maintain stage that starts generating too much noise effectively re-enters tuning. A rule that gets bypassed by a technique variation may need a new hypothesis entirely. Treating this as a cycle rather than a one-way pipeline is what keeps a detection library healthy instead of accumulating dead weight.
Detection-as-Code
The Console-Only Approach
The older way of managing detections looks like this: an analyst notices something suspicious, opens the SIEM console, writes a rule directly in the UI, and saves it. There is no record of who wrote it, why, what it was tested against, or what changed if someone edits it next month. If the rule breaks something or stops working, there is no easy way to see what it looked like before, and nobody outside the person who wrote it fully understands its logic. This works fine for one rule written by one person. It stops working once a team has hundreds of rules maintained by multiple engineers over multiple years.
Applying Software Discipline to Detection Logic
Detection-as-code applies the same discipline to detection logic that software engineering applies to application code.
- Rules live as text files in a version control system, not as opaque objects inside a vendor console.
- Changes go through a review process before merging, the same way a code change would.
- Rules get tested, ideally automatically, before they reach production.
- Every rule carries documentation: what it detects, why it exists, what data source it depends on, and who owns it.
- Because everything is versioned, a bad change can be rolled back to a known-good state instead of being reconstructed from memory.
None of this changes what a detection does. It changes whether the team can trust, audit, and safely evolve a detection library that has grown past what any one person can hold in their head. A team with five rules can probably get away with the console approach. A team with five hundred rules across multiple SIEMs and multiple engineers generally cannot, at least not without accumulating rules nobody remembers writing and nobody is confident is safe to remove.
Where Detections Come From
New detections rarely start from a blank page. They tend to originate from one of a handful of recurring sources, and knowing which source triggered a given detection request often shapes how urgent it is and how it should be prioritized.
- ATT&CK gap analysis: mapping existing detections against the MITRE ATT&CK matrix and identifying techniques, particularly ones relevant to the organization's threat model, that currently have no coverage at all.
- Threat intelligence: a report describing a new or updated technique used by an actor the organization considers relevant, surfaced through a CTI feed, an ISAC bulletin, or vendor research.
- Incident and hunt findings: something that was caught once, whether during an actual incident or during a proactive hunt, that got resolved manually but has no repeatable detection behind it yet. Turning a one-off catch into a standing rule is one of the most common ways new detections get created.
- Vendor and community rule sources: public repositories of Sigma rules and similar community-maintained detection content, which are a reasonable starting point to adapt but not a substitute for tuning to a specific environment's data sources, naming conventions, and normal baseline of activity.
That last point is worth sitting with. A Sigma rule pulled from a public repository was written against someone else's environment and someone else's baseline of what normal looks like. Deploying it unmodified is a starting point, not a finished detection: it still needs to go through the test and tune stages of the lifecycle before it can be trusted in production.
How the Rest of This Module Builds on This Chapter
Everything from here forward drills into one piece of the lifecycle introduced above.
The module moves roughly in the order a detection would actually move through in practice: learning the languages you write rules in, turning a hypothesis into working logic, proving that logic actually works, cutting down the noise it generates, checking it against your broader coverage picture, formalizing how it is managed over time, and finally measuring whether any of this is working.
- Chapter 2: Query Languages: KQL, SPL, and Sigma. The syntax and mental model behind the three languages most detection logic gets written in.
- Chapter 3: From Hypothesis to Detection Logic. How to take a vague idea about attacker behavior and turn it into a scoped, testable rule.
- Chapter 4: Testing and Validation. Proving a rule actually fires on the behavior it's meant to catch, including how Atomic Red Team fits into that process.
- Chapter 5: False Positive Tuning. Narrowing a rule so it survives contact with a real, noisy production environment.
- Chapter 6: Coverage Mapping Against ATT&CK. Using the ATT&CK matrix to see what's actually covered versus what only looks covered.
- Chapter 7: Detection-as-Code and the Rule Lifecycle. The full version-controlled workflow introduced briefly in this chapter, expanded end to end.
- Chapter 8: Metrics and Detection Engineering Maturity. How to know whether a detection program is actually improving, and what to measure to find out.
Key Takeaways
- Detection engineering means turning a hypothesis about attacker behavior into a tuned, production detection rule, distinct from triaging alerts (SOC analyst) or proactively searching for undetected activity (threat hunter).
- The detection lifecycle has six stages: hypothesis, logic, test, tune, deploy, and maintain. It loops rather than running once.
- Detection-as-code treats detection logic like software: version-controlled, peer-reviewed, tested, documented, and reversible. It matters most once a rule library grows past what one person can hold in their head.
- Detections originate from ATT&CK gap analysis, threat intelligence, incident and hunt findings, and community rule sources, each of which needs its own tuning before it can be trusted in production.
- This chapter is the overview. Chapters 2 through 8 each expand one stage of the lifecycle or one supporting skill in depth.
Knowledge Check
Click an answer to reveal the explanation.
A threat hunter finds a suspicious pattern of activity during a manual hunt and confirms it was malicious. What is the most likely next step for a detection engineer?
Which pair correctly matches a role to its primary output?
A team pulls a Sigma rule from a public community repository and deploys it to production without modification. What is the main risk in doing this?