From Hypothesis to Detection Logic
A hypothesis is not a detection. It is a claim about what an attacker's behavior would look like in your environment, and until that claim is turned into a query that runs against real data, it is only an idea. This chapter walks through the mechanical work of getting from a written hypothesis to working logic: picking a data source that can actually see the behavior, choosing the fields that carry signal, deciding how conditions should combine, and mapping the result back to ATT&CK so the detection's purpose stays legible to everyone who reads it later. The goal is a rule you can defend, not just one that runs.
Starting From a Hypothesis, Not a Query
Why Hypothesis Comes Before Query
It's tempting to open a query editor first and start experimenting with fields until something looks interesting. That approach produces detections that are hard to explain, hard to tune, and hard to map to anything. The lifecycle from chapter 1 puts hypothesis before logic for a reason: the hypothesis is what tells you whether the logic you eventually write is even aimed at the right thing.
A detection hypothesis is a specific, falsifiable statement about attacker behavior and how it would surface in available telemetry.
- You can imagine evidence that would prove it wrong, not just evidence that would confirm it.
- "Detect PowerShell abuse" is not falsifiable in that sense; it's a goal, not a claim.
- It doesn't tell you which behavior, which telemetry, or which anomaly you're looking for, so there's no way to know when the resulting rule has succeeded or failed at its actual job.
Vague Goal vs. Well-Formed Hypothesis
| Vague goal | Well-formed hypothesis |
|---|---|
| "Detect PowerShell abuse." | "An adversary using a command and scripting interpreter (T1059) to run obfuscated or encoded commands will spawn as a child process of an unexpected parent, such as an office application or a script host, with a command line containing base64 or explicit execution-policy bypass flags." |
| "Catch brute-force attacks." | "An adversary attempting password guessing (T1110.001) against a single account will produce a burst of failed authentication events for that account from one or a small number of source addresses within a short window, followed in some cases by a success." |
Notice what the well-formed versions have in common: a named behavior, an expected artifact of that behavior, and enough specificity that you could immediately start asking what log source would even contain that artifact. That's the bridge into the next stage. A hypothesis that can't be turned into that kind of question yet usually still needs another pass of refinement before it's ready to leave the whiteboard.
Choosing the Right Data Source
A hypothesis is only as good as the telemetry available to test it. Before writing a single line of logic, ask the plain question: what log source would actually capture this behavior if it happened right now? Different categories of attacker activity live in different data, and picking the wrong one wastes the rest of the lifecycle on a detection that can never fire.
| Behavior category | Likely data source | What it captures |
|---|---|---|
| Process execution, scripting interpreters, LOLBin abuse | Process creation / EDR telemetry | Process name, command line, parent-child lineage, hashes |
| Credential use, logons, privilege changes | Authentication logs (host and identity provider) | Account, logon type, source, success/failure, timestamps |
| Lateral movement, data transfer, C2 beaconing | Network / firewall / proxy logs | Source and destination, ports, bytes, protocol, domain |
| Resource creation, permission changes, API abuse in cloud | Cloud API / control-plane audit logs | Caller identity, action, resource, source IP, outcome |
Is the Source Actually Available
The harder question comes next: is that source actually onboarded, in the SIEM, and reliable in your environment right now? A hypothesis about process-level scripting abuse is dead on arrival if command-line logging isn't enabled, or if the relevant hosts don't forward EDR telemetry yet. This is where detection engineering leans directly on the log-source-onboarding work covered conceptually in the SOC Operations module: onboarding a source isn't a one-time checkbox, it's an ongoing dependency that detection logic sits on top of.
When the Data Doesn't Exist Yet
When the needed source doesn't exist yet, the hypothesis doesn't get discarded, it gets parked. A realistic detection engineering backlog usually has a category for hypotheses that are well-formed but blocked on data availability. Documenting the gap (what source is missing, what it would enable) does two things: it gives the team a concrete ask for the onboarding backlog, and it prevents someone from quietly reinventing the same hypothesis six months later without knowing it was already scoped.
Field Selection and Logic Construction
With a data source confirmed, the next decision is which fields inside it actually carry the signal your hypothesis depends on. Most raw events carry far more fields than any single detection needs, and picking the wrong subset either misses the behavior entirely or buries it in noise.
Which Fields Carry the Signal
- Process-execution hypothesis: process name, full command line, parent process name, the executing user, and sometimes the signing status or file path of the binary.
- Authentication hypothesis: account name, logon type, source address, outcome, and timestamp.
Once the fields are chosen, the logic is mostly about how conditions combine:
- AND narrows a match down to a more specific pattern: a process name matching a scripting interpreter AND a command line containing an encoding flag is far more specific than either condition alone.
- OR covers known variants of the same behavior: several different interpreter binary names, or several known flag spellings, all pointing at the same underlying technique.
- NOT excludes patterns that are technically a match but known to be legitimate in your environment, such as a scripting interpreter launched by a specific, approved deployment tool.
- Thresholds turn a single-event pattern into a frequency-based one, useful when the individual event is unremarkable but the volume or rate is the actual signal, such as the same account failing authentication five or more times inside a two-minute window.
The pseudocode below illustrates the shape of that logic. It isn't tied to any specific query language; the point is the structure, not the syntax.
// pseudocode, not a real query language
IF process.name IN ("interpreter_a.exe", "interpreter_b.exe") // OR: known variants
AND process.command_line CONTAINS ("-enc", "-EncodedCommand") // AND: narrows to encoded execution
AND NOT process.parent.name == "approved_deployment_tool.exe" // NOT: excludes known-legitimate parent
THEN flag_as_candidate
IF COUNT(auth.failure WHERE auth.user == same_user) >= 5
WITHIN 2 MINUTES
THEN flag_as_candidate // threshold: frequency-based, not single-event
Field Selection Is Iterative
Field selection and logic construction feed each other. Sometimes a field you assumed would be reliable turns out to be inconsistently populated once you look at real events, which sends you back to pick a different field or a different combination. That back-and-forth is normal and is part of why this stage of the lifecycle takes longer than it looks like it should on paper.
Single-Event vs. Sequence Detections
Not every hypothesis can be tested against one event in isolation. Some behaviors only make sense as a pattern across time, which splits detection logic into two broad shapes.
Single-Event Detections
A single-event detection evaluates one record against a pattern and fires immediately if it matches. The encoded-command example above is single-event: everything the logic needs (process name, command line, parent) is present in one process-creation record. These are simpler to write, faster to test, and generally cheaper to run, because the SIEM doesn't need to hold state or correlate anything across time.
Sequence Detections
A sequence detection requires two or more related events, often from the same entity, to occur in a particular order within a bounded window. The authentication example extends naturally into one: a burst of failed logons for an account, followed by a success from a source or location that hadn't been seen for that account before, all within a short window.
Neither event alone is necessarily suspicious. A single failed logon happens constantly. A single successful logon from a new location happens too, for ordinary reasons like travel or a new device. It's the combination and the ordering that turns two unremarkable events into a meaningful pattern, which mirrors the kind of correlation logic described at a high level in the SOC Operations module's alert-correlation content.
Sequence detections are more powerful because they capture attacker behavior that genuinely spans multiple steps, but they're harder to build and tune.
- They depend on the SIEM's ability to join or correlate events by a shared entity (account, host, session) across a time window, which not every platform or license tier supports equally well.
- If the window is too short, a real sequence gets missed because the events fall just outside it.
- If it's too long, unrelated events start getting joined together and false positives climb.
Mapping to ATT&CK as You Build
Map Early, Not as an Afterthought
Every detection should map to at least one ATT&CK technique, and ideally the specific sub-technique, from the moment it's written, not bolted on afterward as documentation busywork. Doing the mapping early forces a useful discipline: you have to be able to say, in ATT&CK's own vocabulary, exactly what behavior this rule is meant to catch. If you can't name the technique confidently, that's often a sign the hypothesis itself is still too vague, and it's worth going back a step rather than writing logic against a fuzzy target.
Why This Matters for Coverage Mapping
The payoff shows up later. Chapter 6 covers building a coverage map across your whole detection library, and that exercise only works if every rule already carries its technique and sub-technique tags. Retrofitting ATT&CK mappings onto a library of untagged rules, one by one, after the fact, is slow and error-prone compared to tagging each rule the day it's written. Treat the mapping as part of the rule's metadata, alongside the hypothesis statement and data source, not as a separate documentation pass.
Sub-Techniques Sharpen the Logic
Mapping to a sub-technique specifically, where one exists, also sharpens the logic itself. "Command and Scripting Interpreter" (T1059) covers a wide range of interpreters and behaviors; mapping instead to the specific sub-technique for the interpreter you're actually targeting keeps the rule's scope honest and makes gaps in coverage across sibling sub-techniques visible later.
A Worked Example, Start to Finish
Here's the full path from hypothesis to logic for one common, well-known technique: scheduled-task-based persistence. This is illustrative and generic, not a production rule.
Put in the pseudocode structure from earlier in the chapter, the logic looks like this:
// pseudocode, not a real query language
IF process.name == "scheduling_utility.exe"
AND task.action CONTAINS ("interpreter_a.exe", "interpreter_b.exe", "\\Temp\\")
AND NOT process.user IN (approved_deployment_accounts)
THEN flag_as_candidate // maps to T1053.005
Notice how each earlier stage of the chapter shows up here:
- The hypothesis is specific and falsifiable.
- The data source was checked for availability before logic was written.
- The fields were chosen because they carry the actual signal.
- The AND/NOT combination narrows and excludes deliberately.
- The technique mapping was decided at the same time as the logic, not after.
That's the full path this chapter has been describing, applied to one concrete case.
Key Takeaways
- A detection hypothesis must be specific and falsifiable; a vague goal like "detect PowerShell abuse" can't be tested or judged as succeeded or failed.
- Confirm the data source can actually capture the behavior, and that it's reliably onboarded, before writing any logic; a missing source belongs in the onboarding backlog, not silently ignored.
- Field selection determines whether logic has signal at all; AND narrows, OR covers known variants, NOT excludes known-legitimate patterns, and thresholds convert frequency into a detectable signal.
- Single-event detections fire on one matching record; sequence detections require related events across a time window and depend on the SIEM's correlation capability, making them more powerful but harder to tune.
- Map every detection to an ATT&CK technique, and ideally a sub-technique, at the time it's written, not as an afterthought, so later coverage mapping is possible without retrofitting.
Knowledge Check
Click an answer to reveal the explanation.
Which of the following is the better-formed detection hypothesis?
A hypothesis depends on command-line logging that isn't currently enabled anywhere in the environment. What's the right next step?
Which scenario is best expressed as a sequence detection rather than a single-event detection?