H3AD-REF / REFERENCES / SOC OPERATIONS REFERENCE

Know The Tier.
Trust The Metric.

The three-tier structure most SOCs are built around, Tier 1 triage, Tier 2 investigation and response, Tier 3 threat hunting, paired with the metrics that get quoted in every SOC maturity conversation: MTTD, MTTA, MTTR, dwell time, and the ratios that reveal whether a team is understaffed before anyone says it out loud. Grade the alerts each tier handles against the Admiralty Code reference before they move up the chain.

SOC Analyst Tier Model

There's no single formal spec for this, it's an industry convention taught and applied with the same shape across vendors, MSSPs, and in-house SOCs: three tiers, escalation flows upward as complexity and severity increase, and each tier hands off what it can't confidently close. The exact titles vary by org, the responsibilities and escalation logic below don't.

TierTypical TitleCore ResponsibilitiesEscalates To / When
Tier 1 Triage Analyst / SOC Analyst I First line of defense. Continuously monitors SIEM/EDR/XDR alert queues, follows documented runbooks and playbooks, gathers initial context, and makes the first true-positive/false-positive call. Optimizes for speed and consistency over deep investigation. Anything ambiguous, anything outside the playbook's scope, or anything confirmed malicious goes to Tier 2.
Tier 2 Incident Responder / SOC Analyst II Handles what Tier 1 escalates. Deeper investigation than a playbook covers: correlates events across multiple log sources and tools, scopes the compromise, and executes containment actions (isolate a host, disable an account, block an indicator). Novel or unfamiliar attack techniques, high-severity or multi-system incidents, or anything needing executive/legal coordination goes to Tier 3.
Tier 3 Senior Analyst / Threat Hunter Handles what nothing else could resolve. Advanced investigation and incident leadership, proactive threat hunting, malware and forensic analysis, detection engineering. Often the tier that writes or tunes the playbooks Tier 1 and Tier 2 run on. Nothing above it in the triage chain, findings feed back into new detections and updated playbooks for Tier 1/2 rather than escalating further.
CONTEXT

This Is A Convention, Not A Universal Law

The three-tier model above is the common reference point taught across SOC-structure explainers and used by most MSSPs and enterprise SOCs, but it flexes with org size and maturity. A small in-house SOC may collapse Tier 1 and Tier 2 into one role out of necessity. A large MSSP may split Tier 3 further into dedicated threat-hunting, detection-engineering, and malware-analysis functions. Some orgs also run a SOC Manager/Lead layer above Tier 3, coordination, staffing, reporting, not a triage tier itself, since it doesn't sit in the alert escalation path. Treat the table as the shared vocabulary, not a fixed org chart.

SOC Metrics & KPI Glossary

These terms get thrown around loosely, but they have precise, commonly-agreed definitions, and mixing them up (especially MTTA vs MTTD vs MTTR) misreports where a SOC's actual bottleneck is. Each one measures a different segment of the same timeline: incident starts, gets detected, gets acknowledged, gets resolved.

MetricDefinitionWhy It Matters
MTTD Mean Time to Detect. Average time from when a threat or compromise actually begins to when the SOC identifies it. Total detection time across incidents in a period, divided by incident count. The number that governs dwell time. The slower detection is, the longer an attacker operates unnoticed and the more damage compounds.
MTTA Mean Time to Acknowledge. Average time from when an alert fires to when an analyst actually picks it up and starts working it, the dead time sitting in the queue. Measures queue health and staffing, not investigation skill. A rising MTTA almost always means alert volume is outpacing analyst capacity, not that analysts got slower at their jobs.
MTTR Mean Time to Respond / Resolve. Average time from confirmed detection to containment and resolution. Note: sources aren't fully consistent on the expansion, some use "Respond" (time to confirmed containment), others "Resolve" or "Recovery" (time to full ticket closure). The endpoint shifts by org; the measured segment, action taken after detection, stays the same. Shows how fast the team acts once something is confirmed real. Usually the first number leadership asks for, and the easiest one to game if tracked in isolation (see below).
Dwell Time The total time an attacker was present and active in the environment before detection, effectively the real-world span MTTD is trying to measure and shrink, tracked per incident as an outcome rather than a period-wide average. Directly tied to blast radius. Longer dwell time means more time for lateral movement, privilege escalation, and exfiltration before anyone intervenes.
Alert Volume / Alert-to-Analyst Ratio Number of alerts generated in a period divided by the number of analysts available to triage them. A high ratio predicts burnout, missed alerts, and a rising MTTA before any of those show up as their own metric.
False Positive Rate Percentage of triaged alerts that turn out not to be genuine threats. A chronically high rate signals tuning debt in the detection content itself, not analyst underperformance, and it's the direct driver of alert fatigue.
Escalation Rate Percentage of Tier 1-triaged alerts that get escalated to Tier 2 or Tier 3. Too low can mean real incidents are being closed at Tier 1 that shouldn't be. Too high can mean playbooks under-scope what Tier 1 is actually authorized to resolve on its own.

Where The Tier Model Meets The Metrics

Neither the tier structure nor the KPI glossary means much read in isolation. The metrics only tell you something useful once you know which tier owns the number, and chasing any one of them in isolation creates its own failure mode.

READING THE NUMBERS

A Rising MTTA Is A Staffing Signal, Not A Skill Problem

Where the tier model explains the metric
MTTA lives almost entirely inside Tier 1's queue. When it creeps up, the reflex is to read it as analysts getting slower or less diligent, but paired with the tier model, a rising MTTA more often means alert volume has outpaced the number of people triaging it, or that the playbooks aren't cutting noise fast enough to keep the queue moving. Check alert-to-analyst ratio and false positive rate before treating it as a performance issue.
FAILURE MODE

Optimizing MTTR Alone Rewards Premature Closure

The trap of tracking one metric without its counterweights
Chasing MTTR down in isolation gives an analyst under pressure every reason to mark a ticket resolved before root cause is actually confirmed, closing fast looks identical to closing correctly on a dashboard that only tracks time. Pair MTTR with dwell time and a reopen rate: a fast MTTR next to a rising reopen rate means incidents are being closed, not fixed.
RELATED, NOT THE SAME AXIS
This page answers "which tier handles this and when it escalates." Escalation Matrix & Severity Classification answers a separate question — how bad the incident is and what response SLA it earns. The two compose: severity decides how fast, the tier model decides who.