CHAPTER 03 35 MIN READ INTERMEDIATE

From Hypothesis to Detection Logic

A hypothesis is not a detection. It is a claim about what an attacker's behavior would look like in your environment, and until that claim is turned into a query that runs against real data, it is only an idea. This chapter walks through the mechanical work of getting from a written hypothesis to working logic: picking a data source that can actually see the behavior, choosing the fields that carry signal, deciding how conditions should combine, and mapping the result back to ATT&CK so the detection's purpose stays legible to everyone who reads it later. The goal is a rule you can defend, not just one that runs.

detection logicATT&CK mappingdata sourcesfield selection
Where this fits: Chapter 1 laid out the detection engineering lifecycle as a loop: hypothesis, logic, testing, tuning, deployment, and maintenance. Chapter 2 covered the languages you write that logic in: KQL, SPL, and Sigma, and how each one expresses the same underlying idea differently. This chapter is where those two threads meet. You already know what a hypothesis looks like and what a query looks like; what's missing is the reasoning that connects one to the other. That reasoning, not syntax, is what separates a detection engineer from someone who can write a query.

Starting From a Hypothesis, Not a Query

Why Hypothesis Comes Before Query

It's tempting to open a query editor first and start experimenting with fields until something looks interesting. That approach produces detections that are hard to explain, hard to tune, and hard to map to anything. The lifecycle from chapter 1 puts hypothesis before logic for a reason: the hypothesis is what tells you whether the logic you eventually write is even aimed at the right thing.

A detection hypothesis is a specific, falsifiable statement about attacker behavior and how it would surface in available telemetry.

Falsifiable means:
  • You can imagine evidence that would prove it wrong, not just evidence that would confirm it.
  • "Detect PowerShell abuse" is not falsifiable in that sense; it's a goal, not a claim.
  • It doesn't tell you which behavior, which telemetry, or which anomaly you're looking for, so there's no way to know when the resulting rule has succeeded or failed at its actual job.

Vague Goal vs. Well-Formed Hypothesis

Vague goalWell-formed hypothesis
"Detect PowerShell abuse." "An adversary using a command and scripting interpreter (T1059) to run obfuscated or encoded commands will spawn as a child process of an unexpected parent, such as an office application or a script host, with a command line containing base64 or explicit execution-policy bypass flags."
"Catch brute-force attacks." "An adversary attempting password guessing (T1110.001) against a single account will produce a burst of failed authentication events for that account from one or a small number of source addresses within a short window, followed in some cases by a success."

Notice what the well-formed versions have in common: a named behavior, an expected artifact of that behavior, and enough specificity that you could immediately start asking what log source would even contain that artifact. That's the bridge into the next stage. A hypothesis that can't be turned into that kind of question yet usually still needs another pass of refinement before it's ready to leave the whiteboard.

Note: A hypothesis doesn't need to be perfectly precise on the first write. It needs to be precise enough to test. You'll often revise it once you see what the data actually looks like.

Choosing the Right Data Source

A hypothesis is only as good as the telemetry available to test it. Before writing a single line of logic, ask the plain question: what log source would actually capture this behavior if it happened right now? Different categories of attacker activity live in different data, and picking the wrong one wastes the rest of the lifecycle on a detection that can never fire.

Behavior categoryLikely data sourceWhat it captures
Process execution, scripting interpreters, LOLBin abuseProcess creation / EDR telemetryProcess name, command line, parent-child lineage, hashes
Credential use, logons, privilege changesAuthentication logs (host and identity provider)Account, logon type, source, success/failure, timestamps
Lateral movement, data transfer, C2 beaconingNetwork / firewall / proxy logsSource and destination, ports, bytes, protocol, domain
Resource creation, permission changes, API abuse in cloudCloud API / control-plane audit logsCaller identity, action, resource, source IP, outcome

Is the Source Actually Available

The harder question comes next: is that source actually onboarded, in the SIEM, and reliable in your environment right now? A hypothesis about process-level scripting abuse is dead on arrival if command-line logging isn't enabled, or if the relevant hosts don't forward EDR telemetry yet. This is where detection engineering leans directly on the log-source-onboarding work covered conceptually in the SOC Operations module: onboarding a source isn't a one-time checkbox, it's an ongoing dependency that detection logic sits on top of.

When the Data Doesn't Exist Yet

When the needed source doesn't exist yet, the hypothesis doesn't get discarded, it gets parked. A realistic detection engineering backlog usually has a category for hypotheses that are well-formed but blocked on data availability. Documenting the gap (what source is missing, what it would enable) does two things: it gives the team a concrete ask for the onboarding backlog, and it prevents someone from quietly reinventing the same hypothesis six months later without knowing it was already scoped.

Note: "The data source exists" and "the data source is trustworthy" are separate checks. A source that drops events under load, has inconsistent field names across host versions, or samples instead of logging everything will quietly undermine logic that assumes complete coverage.

Field Selection and Logic Construction

With a data source confirmed, the next decision is which fields inside it actually carry the signal your hypothesis depends on. Most raw events carry far more fields than any single detection needs, and picking the wrong subset either misses the behavior entirely or buries it in noise.

Which Fields Carry the Signal

Typical fields by hypothesis type:
  • Process-execution hypothesis: process name, full command line, parent process name, the executing user, and sometimes the signing status or file path of the binary.
  • Authentication hypothesis: account name, logon type, source address, outcome, and timestamp.

Once the fields are chosen, the logic is mostly about how conditions combine:

  • AND narrows a match down to a more specific pattern: a process name matching a scripting interpreter AND a command line containing an encoding flag is far more specific than either condition alone.
  • OR covers known variants of the same behavior: several different interpreter binary names, or several known flag spellings, all pointing at the same underlying technique.
  • NOT excludes patterns that are technically a match but known to be legitimate in your environment, such as a scripting interpreter launched by a specific, approved deployment tool.
  • Thresholds turn a single-event pattern into a frequency-based one, useful when the individual event is unremarkable but the volume or rate is the actual signal, such as the same account failing authentication five or more times inside a two-minute window.

The pseudocode below illustrates the shape of that logic. It isn't tied to any specific query language; the point is the structure, not the syntax.

// pseudocode, not a real query language

IF process.name IN ("interpreter_a.exe", "interpreter_b.exe")   // OR: known variants
   AND process.command_line CONTAINS ("-enc", "-EncodedCommand")  // AND: narrows to encoded execution
   AND NOT process.parent.name == "approved_deployment_tool.exe"  // NOT: excludes known-legitimate parent
THEN flag_as_candidate

IF COUNT(auth.failure WHERE auth.user == same_user) >= 5
   WITHIN 2 MINUTES
THEN flag_as_candidate   // threshold: frequency-based, not single-event

Field Selection Is Iterative

Field selection and logic construction feed each other. Sometimes a field you assumed would be reliable turns out to be inconsistently populated once you look at real events, which sends you back to pick a different field or a different combination. That back-and-forth is normal and is part of why this stage of the lifecycle takes longer than it looks like it should on paper.

Single-Event vs. Sequence Detections

Not every hypothesis can be tested against one event in isolation. Some behaviors only make sense as a pattern across time, which splits detection logic into two broad shapes.

Single-Event Detections

A single-event detection evaluates one record against a pattern and fires immediately if it matches. The encoded-command example above is single-event: everything the logic needs (process name, command line, parent) is present in one process-creation record. These are simpler to write, faster to test, and generally cheaper to run, because the SIEM doesn't need to hold state or correlate anything across time.

Sequence Detections

A sequence detection requires two or more related events, often from the same entity, to occur in a particular order within a bounded window. The authentication example extends naturally into one: a burst of failed logons for an account, followed by a success from a source or location that hadn't been seen for that account before, all within a short window.

Neither event alone is necessarily suspicious. A single failed logon happens constantly. A single successful logon from a new location happens too, for ordinary reasons like travel or a new device. It's the combination and the ordering that turns two unremarkable events into a meaningful pattern, which mirrors the kind of correlation logic described at a high level in the SOC Operations module's alert-correlation content.

Sequence detections are more powerful because they capture attacker behavior that genuinely spans multiple steps, but they're harder to build and tune.

Why sequences are harder:
  • They depend on the SIEM's ability to join or correlate events by a shared entity (account, host, session) across a time window, which not every platform or license tier supports equally well.
  • If the window is too short, a real sequence gets missed because the events fall just outside it.
  • If it's too long, unrelated events start getting joined together and false positives climb.
Note: When a hypothesis clearly describes a sequence, resist the urge to force it into a single-event rule by only detecting the second half (the success alone, say). That version will fire on plenty of ordinary logons and lose the exact context that made the pattern worth detecting.

Mapping to ATT&CK as You Build

Map Early, Not as an Afterthought

Every detection should map to at least one ATT&CK technique, and ideally the specific sub-technique, from the moment it's written, not bolted on afterward as documentation busywork. Doing the mapping early forces a useful discipline: you have to be able to say, in ATT&CK's own vocabulary, exactly what behavior this rule is meant to catch. If you can't name the technique confidently, that's often a sign the hypothesis itself is still too vague, and it's worth going back a step rather than writing logic against a fuzzy target.

Why This Matters for Coverage Mapping

The payoff shows up later. Chapter 6 covers building a coverage map across your whole detection library, and that exercise only works if every rule already carries its technique and sub-technique tags. Retrofitting ATT&CK mappings onto a library of untagged rules, one by one, after the fact, is slow and error-prone compared to tagging each rule the day it's written. Treat the mapping as part of the rule's metadata, alongside the hypothesis statement and data source, not as a separate documentation pass.

Sub-Techniques Sharpen the Logic

Mapping to a sub-technique specifically, where one exists, also sharpens the logic itself. "Command and Scripting Interpreter" (T1059) covers a wide range of interpreters and behaviors; mapping instead to the specific sub-technique for the interpreter you're actually targeting keeps the rule's scope honest and makes gaps in coverage across sibling sub-techniques visible later.

A Worked Example, Start to Finish

Here's the full path from hypothesis to logic for one common, well-known technique: scheduled-task-based persistence. This is illustrative and generic, not a production rule.

1
Hypothesis
An adversary establishing persistence via scheduled task creation (T1053.005) will create a new scheduled task through a task-scheduling utility, often with an unusual action (a script interpreter or an uncommon binary path) and outside of normal administrative change windows.
→
2
Data source
Process creation telemetry for the task-scheduling utility invocation, plus the corresponding scheduled-task-creation event log if the host generates one. Both need to be confirmed as onboarded before logic is written.
→
3
Field selection
Process name and command-line arguments of the scheduling utility, the task action or command being registered, the executing user, and the parent process that invoked the utility.
→
4
Logic structure
Scheduling utility invocation AND a task action pointing at a script interpreter or an uncommon path, NOT excluding known deployment or patching accounts, with an optional threshold if the same host registers several tasks in quick succession.
→
5
ATT&CK mapping
T1053.005, Scheduled Task/Job: Scheduled Task, under the Persistence and Execution tactics.

Put in the pseudocode structure from earlier in the chapter, the logic looks like this:

// pseudocode, not a real query language

IF process.name == "scheduling_utility.exe"
   AND task.action CONTAINS ("interpreter_a.exe", "interpreter_b.exe", "\\Temp\\")
   AND NOT process.user IN (approved_deployment_accounts)
THEN flag_as_candidate   // maps to T1053.005

Notice how each earlier stage of the chapter shows up here:

  • The hypothesis is specific and falsifiable.
  • The data source was checked for availability before logic was written.
  • The fields were chosen because they carry the actual signal.
  • The AND/NOT combination narrows and excludes deliberately.
  • The technique mapping was decided at the same time as the logic, not after.

That's the full path this chapter has been describing, applied to one concrete case.

Key Takeaways

  • A detection hypothesis must be specific and falsifiable; a vague goal like "detect PowerShell abuse" can't be tested or judged as succeeded or failed.
  • Confirm the data source can actually capture the behavior, and that it's reliably onboarded, before writing any logic; a missing source belongs in the onboarding backlog, not silently ignored.
  • Field selection determines whether logic has signal at all; AND narrows, OR covers known variants, NOT excludes known-legitimate patterns, and thresholds convert frequency into a detectable signal.
  • Single-event detections fire on one matching record; sequence detections require related events across a time window and depend on the SIEM's correlation capability, making them more powerful but harder to tune.
  • Map every detection to an ATT&CK technique, and ideally a sub-technique, at the time it's written, not as an afterthought, so later coverage mapping is possible without retrofitting.

Knowledge Check

Click an answer to reveal the explanation.

Which of the following is the better-formed detection hypothesis?

Option B is falsifiable: it names a specific behavior, an expected parent-child relationship, and a concrete command-line indicator, so you can test whether the pattern actually occurs and judge whether the resulting logic succeeded. The others are goals or general concerns, not testable claims.

A hypothesis depends on command-line logging that isn't currently enabled anywhere in the environment. What's the right next step?

Parking the hypothesis with a clear note of what's missing keeps it usable later and gives the onboarding backlog a concrete, justified request. Deploying logic against data that doesn't exist yet produces a rule that silently never fires, and substituting an unrelated source would test a different behavior than the one hypothesized.

Which scenario is best expressed as a sequence detection rather than a single-event detection?

Option B only becomes meaningful as a combination of two related events in order within a bounded window; neither the failures nor the later success is suspicious on its own. The other options are each fully evaluable from a single record.