Security Operations Basics
Movies show a SOC as a dark room full of screens and someone shouting "we're in." The real thing is quieter and less dramatic: a queue of alerts, a ticketing system, a rotation of analysts working through them one at a time. A SOC is the operational function that watches an organization's environment continuously and decides, alert by alert, whether something needs action. This chapter covers how that function is actually structured, what the tools in front of an analyst really do, and what a shift looks like once the adrenaline of the pitch deck wears off.
SOC Structure: Tier 1, Tier 2, Tier 3
Most SOCs organize analysts into tiers, and the tiers exist for a practical reason: alert volume is high, most alerts are not incidents, and it is inefficient to have your most experienced analyst reading every login notification. The tiered model routes low-complexity work to less experienced staff and reserves deep investigation for people who have earned the pattern recognition to do it fast. It is not a hierarchy of importance so much as a division of labor by depth of analysis required.
Tier 1: The Front Line
Tier 1 analysts handle initial alert review: does this alert make sense given what fired it, is there an obvious explanation, does it match a known benign pattern, and does it need to go further. A Tier 1 analyst working a phishing alert checks headers, checks the sender domain, checks whether other users received the same email, and decides in a few minutes whether this is routine or needs escalation.
Speed and consistency matter more than deep expertise at this tier, because the volume is high and most of what comes through is either a false positive or already-understood benign activity.
Tier 2: Investigation and Containment
Tier 2 analysts pick up what Tier 1 escalates. Their job is deeper investigation: pulling additional telemetry, correlating across multiple log sources, establishing a timeline, and where appropriate taking containment action such as isolating a host or disabling an account.
Tier 2 work requires more context. A Tier 2 analyst investigating a suspicious PowerShell execution needs to understand what normal admin activity looks like on that specific host, whether the parent process is legitimate, and whether the behavior fits a known technique. They also own a lot of the incident coordination work: looping in IT, documenting scope, and deciding when something needs to become a formal incident.
Tier 3: Hunting and Tuning
Tier 3 analysts do the work that does not fit neatly into a ticket queue: proactive threat hunting, advanced malware and forensic analysis, and tuning the detection stack itself. When a Tier 2 investigation surfaces a detection gap, or a new technique shows up in a threat intel report, it is usually Tier 3 that writes the new correlation rule or hunt query.
This tier also tends to own tool administration: SIEM rule libraries, EDR policy tuning, and playbook development that Tier 1 and Tier 2 rely on day to day. Not every SOC has a formal Tier 3; smaller teams often fold this into a senior analyst or detection engineering role.
| Tier | Primary Focus | Typical Actions |
|---|---|---|
| Tier 1 | Initial alert triage | Review, classify, close false positives, escalate real hits |
| Tier 2 | Investigation and containment | Correlate telemetry, build timeline, isolate hosts, coordinate IR |
| Tier 3 | Hunting and tuning | Proactive hunts, malware analysis, detection rule authoring |
What SIEM Actually Does
SIEM stands for Security Information and Event Management. None of what it does sounds dramatic, and it isn't. A SIEM is closer to a very large, well-indexed database with an alerting layer on top than it is to some kind of intelligent detection brain.
- Aggregation: pulls logs from across the environment into one central platform.
- Normalization: maps logs from different vendors into a common schema so fields mean the same thing everywhere.
- Correlation: runs rules against the normalized data to surface events worth an analyst's attention.
Aggregation
Aggregation means pulling logs from firewalls, endpoints, domain controllers, cloud services, proxies, and anything else generating security-relevant events, and shipping them to a central platform. This sounds simple but is the part organizations most often get wrong: a SIEM only sees what is forwarded to it.
If DNS logs are not being shipped, the SIEM cannot alert on suspicious DNS activity no matter how well-written the correlation rule is. Log source coverage is the actual foundation of SOC capability, and gaps in that coverage are invisible until an incident happens and someone asks "do we have logs for that."
Normalization
Normalization takes logs arriving in dozens of different formats, from dozens of different vendors, and maps them into a common schema so a field like "source IP" or "username" means the same thing regardless of which product generated the log. Without normalization, writing a single rule that catches suspicious logins across Windows, a cloud identity provider, and a VPN concentrator would mean writing three separate rules in three separate syntaxes.
Normalization is unglamorous plumbing work, but a SIEM with poor normalization produces correlation rules that quietly miss data because the field they are matching against does not line up the way the analyst assumed.
Correlation Rules
Correlation rules are the logic that turns normalized events into alerts. A rule might say: if a single account fails authentication five times in two minutes and then succeeds, generate an alert. Or: if a process spawns from an Office document and then makes an outbound connection to a rare external domain, generate an alert.
These rules range from simple threshold logic to more complex sequence and behavioral patterns. The quality of a SIEM deployment is measured almost entirely by the quality of the log sources feeding it and the quality of the rules written against those sources, not by which vendor's name is on the license.
This is why the phrase "garbage in, garbage out" gets repeated so often around SIEM deployments. A SIEM with expensive licensing but poor log coverage and untuned rules will generate a flood of low-value alerts while missing the activity that actually matters. A well-scoped deployment with disciplined log source onboarding and continuously tuned rules, even on a smaller platform, will outperform it. The tool is not the detection, the engineering behind the tool is the detection.
What EDR Actually Does
EDR stands for Endpoint Detection and Response, and it does for a single host what a SIEM tries to do across the whole environment: collect detailed telemetry and surface suspicious activity, but with a resolution and response capability that host-level logs alone do not provide. Where a SIEM might see a single Windows event log entry, an EDR agent sees the full process tree, the network connections that process made, the files it touched, and the registry keys it modified, all tied together in a way an analyst can actually pivot through.
Process Tree Visibility
Process tree visibility is the core of what makes EDR useful. Instead of seeing that "powershell.exe ran" in isolation, an analyst sees that Outlook spawned a Word document, Word spawned PowerShell, and PowerShell spawned a network connection to an external IP, all in a single chain with parent-child relationships intact. That chain is what distinguishes a legitimate admin script from a phishing payload executing on a workstation.
Along with process trees, EDR agents track network connections initiated by each process, file creation and modification events, and often registry and driver activity, giving a much richer picture than log files alone.
Detection: Signature-Based vs Behavioral
Detection in EDR platforms happens two ways.
| Method | How It Works | Trade-off |
|---|---|---|
| Signature-based | Matches known-bad indicators, a specific file hash or a known malicious binary | Fast and low-noise, but only catches what has already been seen and cataloged; does nothing against a new or modified payload |
| Behavioral | Watches for patterns regardless of the specific file: process injection into another process's memory, a script disabling security tooling, credential material accessed by a process with no legitimate reason to touch it | Catches novel threats signatures miss, but tends to generate more false positives since legitimate admin tools can look behaviorally similar |
Response Actions
The response side of EDR is what separates it from a pure monitoring tool. These actions compress the time between "we found something" and "the bleeding has stopped" from what used to require physically pulling a network cable to a few clicks. From the console, an analyst can:
- Isolate a host from the network while keeping the EDR agent's management channel alive.
- Kill a running process that is actively executing malicious behavior.
- Quarantine a file so it cannot execute again.
Host isolation in particular is one of the most consequential actions a Tier 2 analyst can take, because it can also disrupt a legitimate business process if the containment call is wrong. That is why isolation decisions usually require some level of confidence or sign-off before they happen.
EDR and SIEM Together
EDR and SIEM are complementary, not competing. EDR gives depth on a single endpoint; SIEM gives breadth across the environment, tying endpoint activity to network, identity, and cloud events. Many organizations forward EDR alerts and telemetry into the SIEM specifically so an analyst can correlate an endpoint detection with a suspicious login from the same account elsewhere in the environment, something neither tool alone would surface.
The Alert Triage Workflow
An alert firing is the start of a process, not an answer. Analysts who skip steps in this sequence, especially under alert volume pressure, are the ones who miss real incidents or waste Tier 2's time on things that could have been resolved at Tier 1.
- True positive: the activity is what the alert says it is and represents a genuine security concern, even if low severity.
- False positive: the underlying activity does not actually match the threat the rule was written to catch.
- Benign true positive: the activity is real and matches the alert, but turns out to be authorized, like a pentester's scan or an approved admin script.
A realistic triage workflow moves through a consistent sequence regardless of what generated the alert:
- Initial review. Read what actually fired: what rule triggered, what raw event caused it, what is the affected asset and user. A surprising number of triage mistakes happen because an analyst pattern-matches on the alert title without reading the underlying event, and closes something that looked like a known benign pattern but was not actually the same activity. Slowing down for thirty seconds at this step saves far more time than it costs.
- Gather context. Pull in everything needed to understand whether the activity is expected: is this user's normal working pattern consistent with the time and location of the alert, is this asset a server that runs scheduled jobs at odd hours or a workstation where 3 a.m. activity is inherently unusual, has this exact alert fired before on this host and what was the disposition then. Good triage pulls from multiple sources: the SIEM, the EDR console, asset inventory, and sometimes a quick check with the user or system owner if time allows.
- Determine true positive vs false positive. This is the analytical core of the job. Watch for benign true positives too, activity that is real and matches the alert but turns out to be authorized, which still needs to be documented rather than silently dismissed.
- Escalate or close. Once the disposition is clear, the analyst either closes the alert with a documented reason or escalates it.
- Document. Closing is not the end of the responsibility. Every closed alert should carry enough documentation that a different analyst, or the same analyst three months later, can understand why it was closed without redoing the investigation from scratch.
Escalation Criteria and Handoffs
Escalation exists so that decisions requiring more context, more authority, or more time get to the person who has them. A Tier 1 analyst who has confirmed activity is malicious but lacks the authority to isolate a production server should escalate immediately rather than sit on the decision. Knowing when to escalate is as much a skill as knowing how to investigate, and it is one of the things that separates a strong Tier 1 analyst from one who either escalates everything or escalates nothing.
When to Escalate
- Confirmed malicious activity. The clearest trigger. Once an analyst has reasonable confidence that what they are looking at is a genuine compromise, whether that is credential theft, malware execution, or unauthorized access, the investigation moves to whoever owns the next stage of response. Escalation criteria are typically written around reasonable confidence, not proof beyond doubt, since waiting for absolute certainty usually costs time the response cannot afford.
- Scope beyond a single host or account. An alert that starts as one suspicious login can turn into a lateral movement pattern touching several systems within minutes. Once an investigation crosses from a single asset to multiple assets, or from a single user to a pattern suggesting a shared credential or a common initial access point, it usually exceeds what a Tier 1 analyst is expected to scope alone and needs Tier 2 involvement to establish the full blast radius.
- Need for containment authority. A practical trigger distinct from technical complexity. Some organizations restrict host isolation, account disablement, or firewall changes to specific roles, regardless of how clear-cut the investigation is, because those actions have business impact and require accountability. A Tier 1 analyst who has done excellent investigative work still needs to hand off to someone with the authority to act if that authority sits outside their role.
What Makes a Handoff Good
What makes a handoff good rather than merely present is specificity. A weak handoff says "this looks suspicious, can you take a look." A strong handoff includes what triggered the alert, what context was gathered, what has already been ruled out, what remains uncertain, and a clear statement of why this is being escalated now rather than closed.
The receiving analyst should be able to pick up the investigation from where it left off rather than start over, which means the handoff needs the same documentation discipline that closing a ticket requires, just aimed at a different reader.
Shift Work and On-Call Realities
Threats do not stop at 5 p.m., so SOCs need coverage that does not either. There are a few common models for achieving this.
Coverage Models
- Follow-the-sun. Uses teams in different time zones, typically three regions roughly eight hours apart, so each team works a normal daytime shift locally while the SOC as a whole covers 24 hours.
- Rotating shifts. Keeps one team but cycles analysts through day, evening, and overnight shifts on a schedule, often rotating every few weeks so no one is stuck permanently on nights.
- Core team plus on-call. Smaller organizations without the headcount for either model above often rely on a smaller core daytime team plus an on-call rotation for after-hours escalations, where someone carries a phone and is expected to respond to a page within a defined window.
The Cost of Off-Hours Work
Overnight and on-call work has a real physiological cost that is worth naming plainly. Working against a normal circadian rhythm degrades attention and decision quality, and rotating shift patterns make it worse than a fixed overnight shift would, because the body never fully adjusts before the schedule changes again. Organizations that ignore this and staff overnight coverage as an afterthought tend to see it show up as slower triage times and more missed context during the hours when fewer senior analysts are watching.
Alert fatigue is the operational problem that shift work makes more visible, though it is not unique to overnight shifts. When an analyst works through a queue of alerts where the overwhelming majority turn out to be false positives, attention and rigor degrade over the course of a shift. A rule that fires accurately once and gets ignored the next fifty times because it fired on benign activity every prior time is a rule that has trained the analyst to distrust it, and that distrust does not discriminate between the fiftieth false positive and the one time it is real. This is a documented and well-studied failure mode in SOC operations, not a personal discipline problem with individual analysts.
Tuning as the Countermeasure
Tuning and automation are the actual countermeasure, and this is why Tier 3 tuning work connects directly back to shift sustainability. Every rule tightened to reduce false positive rate, every repetitive triage step automated through a playbook, and every noisy log source cleaned up reduces the raw volume an analyst has to push through per shift. This is not just an efficiency exercise, it is what keeps the humans doing triage sharp enough to catch the alert that matters. A SOC that treats detection tuning as optional overhead rather than a core operational responsibility will eventually pay for it in a missed incident buried under noise.
None of this means SOC work is grim by default. Well-run SOCs treat shift design, workload balance, and detection tuning as deliberate engineering problems rather than something to solve with more caffeine and longer hours. Analysts who understand these realities going in, and who work somewhere that takes tuning seriously, tend to build sustainable careers in operations rather than burning out in the first two years.
Key Takeaways
- The tiered SOC model routes work by depth: Tier 1 triages, Tier 2 investigates and contains, Tier 3 hunts and tunes the detection stack.
- A SIEM aggregates, normalizes, and correlates logs. It is only as good as the log sources feeding it and the rules written against them.
- EDR gives depth on a single endpoint: process trees, network connections, and file activity, plus response actions like isolation, process kill, and quarantine.
- Triage is a repeatable sequence: initial review, gather context, determine disposition, escalate or close, document. Skipping steps under volume pressure causes missed incidents.
- Escalation triggers on confirmed malicious activity, scope beyond a single host, or a need for containment authority the current tier does not hold. A strong handoff carries full context forward.
- Alert fatigue is a real operational risk, not a discipline failure. Tuning and automation are what protect analyst attention and keep 24/7 coverage sustainable.
Knowledge Check
Click an answer to reveal the explanation.
A Tier 1 analyst reviews an alert, confirms it matches an approved IT change ticket, and closes it with the note "not malicious." What is the problem with this handling?
During investigation of a suspicious login, a Tier 1 analyst discovers the same credential was used to authenticate successfully on three additional servers within the last ten minutes. What should happen next?
A SIEM correlation rule for suspicious authentication has fired hundreds of times over the past month, and every instance investigated so far has been a benign VPN reconnect pattern. What is the most appropriate response?