SIEM and Log Management
Chapter 2 covered what to do once an alert lands in your queue: how to weigh severity, asset value, and context into a triage decision. This chapter goes one layer deeper and asks where that alert actually came from. A SIEM is the machine that turns raw log lines scattered across hundreds of systems into the single correlated alert you triaged. Understanding its pipeline, from collection through correlation, changes how you read an alert: you start seeing the data behind it instead of just the label on top of it.
What a SIEM Actually Does
A SIEM (Security Information and Event Management platform) is often described as a single tool, but it is better understood as a pipeline with four distinct stages. An analyst who only ever sees the search bar or the alert dashboard is seeing the last stage of a process that started much earlier, on an endpoint, a firewall, or a domain controller, and passed through several transformations before it became something searchable.
| Stage | What Happens | Example |
|---|---|---|
| Collection | Raw events are pulled or pushed from source systems using agents, forwarders, or API integrations, then shipped to the SIEM | A lightweight forwarder installed on a Windows server ships Security event log entries to the SIEM's ingestion layer |
| Normalization | Raw, source-specific formats are parsed into a common schema so a "source IP" field means the same thing whether it came from a firewall or a proxy | A firewall's proprietary syslog format and a cloud provider's JSON audit log both get mapped to shared fields like src_ip, user, and event_action |
| Correlation | Rules or analytics scan normalized events, often across multiple sources, looking for patterns that indicate something worth an analyst's attention | A rule fires when the same account fails authentication five times within a minute and then succeeds from a new location |
| Storage and Search | Normalized events are indexed and retained so analysts can query historical data, build dashboards, and investigate beyond the immediate alert | An analyst searches thirty days of proxy logs for every connection to a domain flagged in a new threat intel report |
Collection Methods
How an event gets from its source to the SIEM varies by platform and by what's generating the data.
| Method | Details |
|---|---|
| Agents | Installed on endpoints, tail local logs, and forward them in near real time. They tend to give you the richest, most granular telemetry because the agent runs on the box itself. |
| Forwarders | Sit closer to infrastructure that can't run an agent, a network appliance or a legacy system, and relay syslog or similar output. |
| API-based collection | Pulls data from cloud platforms and SaaS applications that expose their logs through a management API rather than a traditional log file. This is now the dominant pattern for cloud control-plane and identity provider logging. |
Log Sources and What to Onboard First
No SOC onboards every available log source on day one. Ingestion has a real cost, in licensing, storage, and the engineering time to parse and normalize each new source correctly, so onboarding is sequenced. The general principle: prioritize the sources that cover the areas an attacker is most likely to touch early and often, then layer in lower-signal sources as maturity and budget allow.
| Category | Typical Sources | What It Reveals |
|---|---|---|
| Endpoint | EDR telemetry, OS-level event logs (process creation, file writes, registry changes) | What actually ran on a host: process lineage, command lines, persistence mechanisms, malware behavior |
| Network | Firewall, web proxy, DNS resolver logs | What talked to what, over which ports, to which domains, and whether traffic matches known bad infrastructure |
| Identity | Directory service logs, SSO and identity provider (IdP) sign-in logs | Who authenticated, from where, with what method, and whether that pattern matches the account's normal behavior |
| Cloud | Control-plane and audit logs from cloud platforms and SaaS admin consoles | Configuration changes, privilege grants, resource creation and deletion, API calls made outside the console |
Why Identity and Endpoint Usually Come First
Most intrusions eventually need a valid credential and a place to execute code, which is why identity and endpoint logging tend to sit at the top of onboarding priority lists. A stolen credential used to sign into an IdP, or a malicious script executed on a workstation, produces a log entry in exactly these two categories before it produces one anywhere else.
Network logs remain valuable, they're often the only place lateral movement or command-and-control beaconing shows up cleanly, but they answer "where did traffic go" rather than "who did it and what did they run," which is usually the more urgent question during an active investigation.
Cloud control-plane logs have become a higher onboarding priority than they were a few years ago, simply because more of an organization's actual infrastructure now lives behind a cloud console rather than a physical firewall. An attacker who compromises a cloud administrator account can do more damage through the API than through anything a network sensor would ever see.
Correlation Rules vs. Behavioral Analytics (UEBA)
Once data is flowing and normalized, the SIEM needs a way to decide which events matter enough to surface as an alert. Two broad approaches dominate: static correlation rules and behavioral analytics, often marketed as UEBA (User and Entity Behavior Analytics). They aren't competitors so much as complements, each catching a different shape of threat.
| Correlation Rules | Behavioral Analytics (UEBA) | |
|---|---|---|
| How it works | Matches a known pattern: specific event sequences, thresholds, or field combinations defined ahead of time | Builds a baseline of normal behavior per user or entity, then flags statistically significant deviations from that baseline |
| Strength | Precise, explainable, fast to write for a known technique. An analyst can read the rule logic and know exactly why it fired | Can surface behavior nobody wrote a rule for: a legitimate account suddenly accessing data it never touches, at hours it never works |
| Blind spot | Only catches what someone thought to write a rule for. A technique with no matching rule produces silence, not a low-confidence alert | Needs enough historical data to establish a meaningful baseline, and a slow, patient attacker who looks "normal enough" can stay under the deviation threshold |
| Tuning burden | Every new rule needs manual tuning against your environment's normal noise | Baselines need to be periodically revalidated, especially after role changes, reorganizations, or new tooling that shifts what "normal" looks like |
In practice, correlation rules remain the backbone of most detection programs because they're explainable and cheap to reason about during an incident, while behavioral analytics fills the gap for slow, low-and-slow activity or insider-style abuse of legitimate access that no static rule was ever written to catch. Neither approach replaces the other; a mature detection program layers both.
The Detection Content Lifecycle
A correlation rule isn't written once and left alone. Detection content has a lifecycle, and skipping stages in that lifecycle is one of the most common reasons a SIEM accumulates rules that are either noisy enough to be ignored or so narrowly tuned they no longer catch anything.
The step teams skip most often is the last one. A rule written two years ago for an environment that has since changed, a decommissioned application, a retired authentication method, a tool nobody uses anymore, can keep firing on noise long after it stopped being useful. Reviewing detection content on a regular cadence and retiring what no longer earns its keep is as much a part of the lifecycle as writing the rule in the first place.
Retention and Compliance, Briefly
Retention Follows Compliance, Not Preference
How long log data stays searchable in a SIEM is rarely a purely technical decision. Retention periods are frequently set by regulatory or contractual obligations that apply to the organization's industry or customer agreements, and the SOC's job is usually to meet whatever floor those obligations set rather than to choose an arbitrary number independently.
If your organization is subject to specific compliance frameworks, the retention requirement should come from whoever owns that compliance relationship, not from an assumption made at the SIEM console.
Hot and Cold Storage Tiers
Underneath the compliance question sits a cost question that shapes day-to-day architecture. Log volume at scale gets expensive to keep fully indexed and instantly searchable, so most SIEM deployments split retention into tiers.
| Tier | What It Means |
|---|---|
| Hot | Data is fully indexed and fast to query |
| Cold / archived | Data is retained to satisfy retention obligations but takes longer, and sometimes a manual restoration step, to search |
Where that line falls between hot and cold is a tradeoff every SOC has to make deliberately, because setting it purely by cost without regard to investigation needs means an analyst three months into an incident may find the logs they need have already aged out of fast search.
Common SIEM Platforms
Most analysts will work with one or more of a small handful of platforms over the course of a career. Each has its own query language, deployment model, and ecosystem, but the underlying pipeline described earlier in this chapter, collect, normalize, correlate, store, applies to all of them.
| Platform | Query Language | Notable Characteristics |
|---|---|---|
| Splunk | SPL (Search Processing Language) | Widely used across enterprise SOCs, deployable on-premises or in the cloud, extensive app and add-on ecosystem for parsing specific log sources |
| Microsoft Sentinel | KQL (Kusto Query Language) | Cloud-native, built on Azure, integrates closely with Microsoft's own identity and endpoint telemetry as well as third-party connectors |
| IBM QRadar | AQL (Ariel Query Language), plus a rule-based correlation engine | Long-established in enterprise and government environments, strong built-in offense/correlation model out of the box |
| Elastic Security | KQL / EQL / Lucene, built on the Elastic (ELK) stack | Open architecture built on Elasticsearch, popular where teams want more control over indexing and storage design or already run the Elastic stack |
The platform matters less than the fundamentals in this chapter. An analyst who understands collection, normalization, correlation, and the detection content lifecycle can move between Splunk, Sentinel, QRadar, or Elastic Security and be productive within days, because the query syntax is the part that changes, not the underlying logic of what a SIEM is trying to do.
Key Takeaways
- A SIEM is a pipeline, not a single tool: collection, normalization, correlation, and storage/search are distinct stages, and a breakdown at any one of them changes what alerts you see.
- Onboarding priority generally follows what attackers touch most: identity and endpoint logs typically come before lower-signal network and cloud sources, though the right order depends on your specific environment.
- Correlation rules and behavioral analytics (UEBA) catch different shapes of threat. Rules are explainable but only catch what someone wrote a rule for; behavioral analytics can surface the unknown but needs a reliable baseline and can miss patient, low-and-slow activity.
- Detection content has a lifecycle: idea, draft, test against historical data, tune, deploy, and periodically review or retire. Skipping the review/retire stage is a common source of stale, noisy rules.
- Retention periods are usually set by regulatory or contractual requirements, and cost tradeoffs shape what stays in fast, hot search versus what moves to slower archive tiers.
- Splunk, Microsoft Sentinel, IBM QRadar, and Elastic Security differ in query language and deployment model, but all implement the same underlying collect/normalize/correlate/store pipeline.
Knowledge Check
Click an answer to reveal the explanation.
Q1. Which SIEM pipeline stage is responsible for making sure a "source IP" field means the same thing whether the event came from a firewall or a cloud audit log?
Q2. Why do identity and endpoint logs typically get onboarded before lower-signal sources?
Q3. What is the main blind spot of static correlation rules compared to behavioral analytics (UEBA)?