Know The Tier.
Trust The Metric.
The three-tier structure most SOCs are built around, Tier 1 triage, Tier 2 investigation and response, Tier 3 threat hunting, paired with the metrics that get quoted in every SOC maturity conversation: MTTD, MTTA, MTTR, dwell time, and the ratios that reveal whether a team is understaffed before anyone says it out loud. Grade the alerts each tier handles against the Admiralty Code reference before they move up the chain.
SOC Analyst Tier Model
There's no single formal spec for this, it's an industry convention taught and applied with the same shape across vendors, MSSPs, and in-house SOCs: three tiers, escalation flows upward as complexity and severity increase, and each tier hands off what it can't confidently close. The exact titles vary by org, the responsibilities and escalation logic below don't.
| Tier | Typical Title | Core Responsibilities | Escalates To / When |
|---|---|---|---|
| Tier 1 | Triage Analyst / SOC Analyst I | First line of defense. Continuously monitors SIEM/EDR/XDR alert queues, follows documented runbooks and playbooks, gathers initial context, and makes the first true-positive/false-positive call. Optimizes for speed and consistency over deep investigation. | Anything ambiguous, anything outside the playbook's scope, or anything confirmed malicious goes to Tier 2. |
| Tier 2 | Incident Responder / SOC Analyst II | Handles what Tier 1 escalates. Deeper investigation than a playbook covers: correlates events across multiple log sources and tools, scopes the compromise, and executes containment actions (isolate a host, disable an account, block an indicator). | Novel or unfamiliar attack techniques, high-severity or multi-system incidents, or anything needing executive/legal coordination goes to Tier 3. |
| Tier 3 | Senior Analyst / Threat Hunter | Handles what nothing else could resolve. Advanced investigation and incident leadership, proactive threat hunting, malware and forensic analysis, detection engineering. Often the tier that writes or tunes the playbooks Tier 1 and Tier 2 run on. | Nothing above it in the triage chain, findings feed back into new detections and updated playbooks for Tier 1/2 rather than escalating further. |
This Is A Convention, Not A Universal Law
SOC Metrics & KPI Glossary
These terms get thrown around loosely, but they have precise, commonly-agreed definitions, and mixing them up (especially MTTA vs MTTD vs MTTR) misreports where a SOC's actual bottleneck is. Each one measures a different segment of the same timeline: incident starts, gets detected, gets acknowledged, gets resolved.
| Metric | Definition | Why It Matters |
|---|---|---|
| MTTD | Mean Time to Detect. Average time from when a threat or compromise actually begins to when the SOC identifies it. Total detection time across incidents in a period, divided by incident count. | The number that governs dwell time. The slower detection is, the longer an attacker operates unnoticed and the more damage compounds. |
| MTTA | Mean Time to Acknowledge. Average time from when an alert fires to when an analyst actually picks it up and starts working it, the dead time sitting in the queue. | Measures queue health and staffing, not investigation skill. A rising MTTA almost always means alert volume is outpacing analyst capacity, not that analysts got slower at their jobs. |
| MTTR | Mean Time to Respond / Resolve. Average time from confirmed detection to containment and resolution. Note: sources aren't fully consistent on the expansion, some use "Respond" (time to confirmed containment), others "Resolve" or "Recovery" (time to full ticket closure). The endpoint shifts by org; the measured segment, action taken after detection, stays the same. | Shows how fast the team acts once something is confirmed real. Usually the first number leadership asks for, and the easiest one to game if tracked in isolation (see below). |
| Dwell Time | The total time an attacker was present and active in the environment before detection, effectively the real-world span MTTD is trying to measure and shrink, tracked per incident as an outcome rather than a period-wide average. | Directly tied to blast radius. Longer dwell time means more time for lateral movement, privilege escalation, and exfiltration before anyone intervenes. |
| Alert Volume / Alert-to-Analyst Ratio | Number of alerts generated in a period divided by the number of analysts available to triage them. | A high ratio predicts burnout, missed alerts, and a rising MTTA before any of those show up as their own metric. |
| False Positive Rate | Percentage of triaged alerts that turn out not to be genuine threats. | A chronically high rate signals tuning debt in the detection content itself, not analyst underperformance, and it's the direct driver of alert fatigue. |
| Escalation Rate | Percentage of Tier 1-triaged alerts that get escalated to Tier 2 or Tier 3. | Too low can mean real incidents are being closed at Tier 1 that shouldn't be. Too high can mean playbooks under-scope what Tier 1 is actually authorized to resolve on its own. |
Where The Tier Model Meets The Metrics
Neither the tier structure nor the KPI glossary means much read in isolation. The metrics only tell you something useful once you know which tier owns the number, and chasing any one of them in isolation creates its own failure mode.