CHAPTER 03 35 MIN READ INTERMEDIATE

SIEM and Log Management

Chapter 2 covered what to do once an alert lands in your queue: how to weigh severity, asset value, and context into a triage decision. This chapter goes one layer deeper and asks where that alert actually came from. A SIEM is the machine that turns raw log lines scattered across hundreds of systems into the single correlated alert you triaged. Understanding its pipeline, from collection through correlation, changes how you read an alert: you start seeing the data behind it instead of just the label on top of it.

SIEM log management correlation rules log sources
Before you continue: This chapter assumes you're comfortable with the triage workflow from Chapter 2, prioritizing an alert queue, reading severity against asset context, and escalation thresholds. Here the question shifts upstream: what pipeline produced that alert in the first place, and why does one log source get onboarded before another?

What a SIEM Actually Does

A SIEM (Security Information and Event Management platform) is often described as a single tool, but it is better understood as a pipeline with four distinct stages. An analyst who only ever sees the search bar or the alert dashboard is seeing the last stage of a process that started much earlier, on an endpoint, a firewall, or a domain controller, and passed through several transformations before it became something searchable.

Stage What Happens Example
Collection Raw events are pulled or pushed from source systems using agents, forwarders, or API integrations, then shipped to the SIEM A lightweight forwarder installed on a Windows server ships Security event log entries to the SIEM's ingestion layer
Normalization Raw, source-specific formats are parsed into a common schema so a "source IP" field means the same thing whether it came from a firewall or a proxy A firewall's proprietary syslog format and a cloud provider's JSON audit log both get mapped to shared fields like src_ip, user, and event_action
Correlation Rules or analytics scan normalized events, often across multiple sources, looking for patterns that indicate something worth an analyst's attention A rule fires when the same account fails authentication five times within a minute and then succeeds from a new location
Storage and Search Normalized events are indexed and retained so analysts can query historical data, build dashboards, and investigate beyond the immediate alert An analyst searches thirty days of proxy logs for every connection to a domain flagged in a new threat intel report

Collection Methods

How an event gets from its source to the SIEM varies by platform and by what's generating the data.

Method Details
Agents Installed on endpoints, tail local logs, and forward them in near real time. They tend to give you the richest, most granular telemetry because the agent runs on the box itself.
Forwarders Sit closer to infrastructure that can't run an agent, a network appliance or a legacy system, and relay syslog or similar output.
API-based collection Pulls data from cloud platforms and SaaS applications that expose their logs through a management API rather than a traditional log file. This is now the dominant pattern for cloud control-plane and identity provider logging.
Note: Normalization is the unglamorous stage that makes everything downstream possible. A correlation rule that says "flag failed logon followed by successful logon from a new IP within five minutes" only works if failed and successful logons from every log source in the environment already share the same field names and value formats. Poorly normalized data quietly breaks rules without ever throwing an error, which is why a rule that used to fire can go silent after an unrelated log source change.

Log Sources and What to Onboard First

No SOC onboards every available log source on day one. Ingestion has a real cost, in licensing, storage, and the engineering time to parse and normalize each new source correctly, so onboarding is sequenced. The general principle: prioritize the sources that cover the areas an attacker is most likely to touch early and often, then layer in lower-signal sources as maturity and budget allow.

Category Typical Sources What It Reveals
Endpoint EDR telemetry, OS-level event logs (process creation, file writes, registry changes) What actually ran on a host: process lineage, command lines, persistence mechanisms, malware behavior
Network Firewall, web proxy, DNS resolver logs What talked to what, over which ports, to which domains, and whether traffic matches known bad infrastructure
Identity Directory service logs, SSO and identity provider (IdP) sign-in logs Who authenticated, from where, with what method, and whether that pattern matches the account's normal behavior
Cloud Control-plane and audit logs from cloud platforms and SaaS admin consoles Configuration changes, privilege grants, resource creation and deletion, API calls made outside the console

Why Identity and Endpoint Usually Come First

Most intrusions eventually need a valid credential and a place to execute code, which is why identity and endpoint logging tend to sit at the top of onboarding priority lists. A stolen credential used to sign into an IdP, or a malicious script executed on a workstation, produces a log entry in exactly these two categories before it produces one anywhere else.

Network logs remain valuable, they're often the only place lateral movement or command-and-control beaconing shows up cleanly, but they answer "where did traffic go" rather than "who did it and what did they run," which is usually the more urgent question during an active investigation.

Cloud control-plane logs have become a higher onboarding priority than they were a few years ago, simply because more of an organization's actual infrastructure now lives behind a cloud console rather than a physical firewall. An attacker who compromises a cloud administrator account can do more damage through the API than through anything a network sensor would ever see.

Note: Onboarding order isn't a fixed checklist. A SOC for a software company with almost no on-premises footprint will prioritize cloud and identity logs far above network logs, while a SOC for a manufacturing environment with a large industrial network may weight network and endpoint sources differently. The "attacker touches it most" heuristic is a starting point, not a substitute for knowing your own environment.

Correlation Rules vs. Behavioral Analytics (UEBA)

Once data is flowing and normalized, the SIEM needs a way to decide which events matter enough to surface as an alert. Two broad approaches dominate: static correlation rules and behavioral analytics, often marketed as UEBA (User and Entity Behavior Analytics). They aren't competitors so much as complements, each catching a different shape of threat.

Correlation Rules Behavioral Analytics (UEBA)
How it works Matches a known pattern: specific event sequences, thresholds, or field combinations defined ahead of time Builds a baseline of normal behavior per user or entity, then flags statistically significant deviations from that baseline
Strength Precise, explainable, fast to write for a known technique. An analyst can read the rule logic and know exactly why it fired Can surface behavior nobody wrote a rule for: a legitimate account suddenly accessing data it never touches, at hours it never works
Blind spot Only catches what someone thought to write a rule for. A technique with no matching rule produces silence, not a low-confidence alert Needs enough historical data to establish a meaningful baseline, and a slow, patient attacker who looks "normal enough" can stay under the deviation threshold
Tuning burden Every new rule needs manual tuning against your environment's normal noise Baselines need to be periodically revalidated, especially after role changes, reorganizations, or new tooling that shifts what "normal" looks like

In practice, correlation rules remain the backbone of most detection programs because they're explainable and cheap to reason about during an incident, while behavioral analytics fills the gap for slow, low-and-slow activity or insider-style abuse of legitimate access that no static rule was ever written to catch. Neither approach replaces the other; a mature detection program layers both.

Tip: When you're reading an alert, note which mechanism generated it. A correlation-rule alert tells you exactly which pattern matched, which is useful for scoping your investigation. A UEBA alert tells you something deviated from baseline, but the "why" is still yours to determine, which usually means the investigation starts a step further back.

The Detection Content Lifecycle

A correlation rule isn't written once and left alone. Detection content has a lifecycle, and skipping stages in that lifecycle is one of the most common reasons a SIEM accumulates rules that are either noisy enough to be ignored or so narrowly tuned they no longer catch anything.

1
Idea / Hypothesis
A gap in coverage, a threat intel report, or a hunt finding suggests a detectable pattern
→
2
Draft Rule
Translate the pattern into query logic against the fields the SIEM actually has
→
3
Test Against Historical Data
Run the draft against past logs to see how often it would have fired and on what
→
4
Tune for False Positives
Add exclusions or refine thresholds against known-legitimate activity surfaced in testing
→
5
Deploy to Production
Rule goes live, routed to the appropriate queue with the appropriate severity
→
6
Review / Retire
Periodically revisit: is it still firing accurately, still relevant, still worth the analyst attention it consumes

The step teams skip most often is the last one. A rule written two years ago for an environment that has since changed, a decommissioned application, a retired authentication method, a tool nobody uses anymore, can keep firing on noise long after it stopped being useful. Reviewing detection content on a regular cadence and retiring what no longer earns its keep is as much a part of the lifecycle as writing the rule in the first place.

Note: Testing against historical data before deployment matters more than it sounds. A rule that looks reasonable on paper can generate hundreds of alerts a day once run against real traffic, usually because some legitimate, high-volume process happens to match the pattern. Catching that in testing is far cheaper than catching it after the rule has buried an analyst queue in false positives for a week.

Retention and Compliance, Briefly

Retention Follows Compliance, Not Preference

How long log data stays searchable in a SIEM is rarely a purely technical decision. Retention periods are frequently set by regulatory or contractual obligations that apply to the organization's industry or customer agreements, and the SOC's job is usually to meet whatever floor those obligations set rather than to choose an arbitrary number independently.

If your organization is subject to specific compliance frameworks, the retention requirement should come from whoever owns that compliance relationship, not from an assumption made at the SIEM console.

Hot and Cold Storage Tiers

Underneath the compliance question sits a cost question that shapes day-to-day architecture. Log volume at scale gets expensive to keep fully indexed and instantly searchable, so most SIEM deployments split retention into tiers.

Tier What It Means
Hot Data is fully indexed and fast to query
Cold / archived Data is retained to satisfy retention obligations but takes longer, and sometimes a manual restoration step, to search

Where that line falls between hot and cold is a tradeoff every SOC has to make deliberately, because setting it purely by cost without regard to investigation needs means an analyst three months into an incident may find the logs they need have already aged out of fast search.

Note: Retention and detection are related but separate decisions. A log source can be fully onboarded for real-time correlation while still following its own retention schedule for historical search. Know both numbers for your critical sources: how long they're actively monitored, and how long they're actually retrievable if an investigation needs to look back further than the hot tier covers.

Common SIEM Platforms

Most analysts will work with one or more of a small handful of platforms over the course of a career. Each has its own query language, deployment model, and ecosystem, but the underlying pipeline described earlier in this chapter, collect, normalize, correlate, store, applies to all of them.

Platform Query Language Notable Characteristics
Splunk SPL (Search Processing Language) Widely used across enterprise SOCs, deployable on-premises or in the cloud, extensive app and add-on ecosystem for parsing specific log sources
Microsoft Sentinel KQL (Kusto Query Language) Cloud-native, built on Azure, integrates closely with Microsoft's own identity and endpoint telemetry as well as third-party connectors
IBM QRadar AQL (Ariel Query Language), plus a rule-based correlation engine Long-established in enterprise and government environments, strong built-in offense/correlation model out of the box
Elastic Security KQL / EQL / Lucene, built on the Elastic (ELK) stack Open architecture built on Elasticsearch, popular where teams want more control over indexing and storage design or already run the Elastic stack

The platform matters less than the fundamentals in this chapter. An analyst who understands collection, normalization, correlation, and the detection content lifecycle can move between Splunk, Sentinel, QRadar, or Elastic Security and be productive within days, because the query syntax is the part that changes, not the underlying logic of what a SIEM is trying to do.

Key Takeaways

  • A SIEM is a pipeline, not a single tool: collection, normalization, correlation, and storage/search are distinct stages, and a breakdown at any one of them changes what alerts you see.
  • Onboarding priority generally follows what attackers touch most: identity and endpoint logs typically come before lower-signal network and cloud sources, though the right order depends on your specific environment.
  • Correlation rules and behavioral analytics (UEBA) catch different shapes of threat. Rules are explainable but only catch what someone wrote a rule for; behavioral analytics can surface the unknown but needs a reliable baseline and can miss patient, low-and-slow activity.
  • Detection content has a lifecycle: idea, draft, test against historical data, tune, deploy, and periodically review or retire. Skipping the review/retire stage is a common source of stale, noisy rules.
  • Retention periods are usually set by regulatory or contractual requirements, and cost tradeoffs shape what stays in fast, hot search versus what moves to slower archive tiers.
  • Splunk, Microsoft Sentinel, IBM QRadar, and Elastic Security differ in query language and deployment model, but all implement the same underlying collect/normalize/correlate/store pipeline.

Knowledge Check

Click an answer to reveal the explanation.

Q1. Which SIEM pipeline stage is responsible for making sure a "source IP" field means the same thing whether the event came from a firewall or a cloud audit log?

B is correct. Normalization parses source-specific raw formats into a common schema so fields align across every log source, which is what makes cross-source correlation rules possible in the first place. Collection (A) is just getting the raw event to the SIEM. Correlation (C) runs after normalization and depends on it. Storage and search (D) is where normalized data gets indexed and retained.

Q2. Why do identity and endpoint logs typically get onboarded before lower-signal sources?

B is correct. A stolen credential used against an identity provider and malicious code executed on a workstation both generate log entries in exactly these two categories before they show up anywhere else, which is why onboarding priority usually follows "what an attacker touches most." Cost (A) is a real factor but not the driver of this specific priority order. Regulatory mandates (C) vary by industry and aren't a universal rule. Normalization difficulty (D) isn't the reason for onboarding sequencing.

Q3. What is the main blind spot of static correlation rules compared to behavioral analytics (UEBA)?

B is correct. Correlation rules match known patterns defined ahead of time, so anything outside that known pattern space generates no alert at all rather than a weaker one. This is precisely the gap behavioral analytics is meant to fill, by flagging statistically significant deviation from a baseline even when no one wrote a rule for the specific technique. Rules are actually more explainable, not less (A is backwards). It's UEBA, not correlation rules, that needs historical data to build a baseline (C is backwards). Correlation rules commonly do span multiple log sources (D is false).