Metrics and KPIs
A SOC that cannot answer "how are we doing" in concrete terms is running on instinct, and instinct doesn't hold up well against a budget review or an auditor's questions. Metrics turn the day-to-day work of triage, escalation, and response into numbers that leadership, the SOC itself, and outside reviewers can actually look at. This chapter covers the core time-based metrics analysts hear most often, the coverage metrics that tell a SOC what it can and can't see, the maturity models that assess a program beyond any single number, and the ways metrics programs quietly go wrong when the numbers get chased instead of the outcomes behind them. None of this is about picking a "correct" set of KPIs; it's about understanding what each metric actually measures, so it gets used honestly instead of gamed or misread.
Why Metrics Matter to a SOC
A SOC produces a huge amount of activity: alerts triaged, cases opened and closed, escalations sent, incidents handled. None of that activity means anything to someone outside the SOC unless it's translated into a small set of numbers that answer a specific question for a specific audience. Metrics exist to do that translation, and different audiences need different translations of the same underlying work.
Leadership
Leadership doesn't see the queue, doesn't see the individual alerts, and generally doesn't want to. What leadership needs is visibility into whether the SOC is actually effective, and whether the resourcing it has is enough to keep it that way. A trend line showing detection and response times holding steady while alert volume climbs is a much stronger staffing argument than an analyst simply saying the team feels stretched.
Metrics give leadership a defensible basis for decisions about headcount, tooling budget, and where the SOC sits in the organization's priorities, decisions that get made whether or not the SOC hands over good numbers to inform them.
The SOC Itself
Metrics aren't only for people outside the SOC. Internally, they function as a tuning feedback loop. A rising average time-to-triage might point to alert fatigue from a noisy detection rule that needs retuning. A spike in reopened tickets might mean an escalation criterion is too loose, or that L1 analysts need more guidance on a particular alert type.
Without metrics, these problems get noticed anecdotally, if at all, usually after they've already cost the team real time. With metrics tracked over time, a SOC lead can catch a drifting trend before it becomes a pattern everyone has quietly gotten used to.
Compliance and Audit
Many organizations operate under a framework, contractual obligation, or regulatory requirement that expects a security operations function to exist and to actually work, not just exist on paper. Metrics are the evidence that backs up that claim.
An auditor reviewing a SOC's incident response process wants to see that alerts get triaged within a defined window, that incidents get documented, and that the numbers behind those claims are tracked consistently rather than reconstructed after the fact. Metrics turn "we have a process" into something a third party can independently verify.
Core Time-Based Metrics
Time-based metrics are the most commonly cited numbers in SOC reporting, and also the most commonly misused, largely because the acronyms sound more standardized than they actually are. Different organizations, and even different vendors, define the exact start and end points of these measurements slightly differently.
The practical approach is to understand what each metric measures conceptually and confirm how your own organization defines its boundaries, rather than assuming the acronym means the same thing everywhere.
| Metric | What It Measures | Why It Matters |
|---|---|---|
| MTTD (Mean Time to Detect) | The average time between when malicious or suspicious activity actually begins and when the SOC (or its tooling) first identifies it as something worth attention | A long MTTD means an attacker has more time to move laterally, escalate privileges, or exfiltrate data before anyone notices. It's a direct reflection of detection coverage and log visibility. |
| MTTA (Mean Time to Acknowledge) | The average time between an alert firing and a human analyst actually picking it up to begin triage | A long MTTA usually points to queue backlog, alert fatigue, or staffing gaps rather than a detection problem. It's the clearest signal of whether alerts are being worked in a timely way once they exist. |
| MTTR (Mean Time to Respond / Remediate) | The average time between an alert being acknowledged and the incident being fully contained, remediated, or otherwise closed out. Some organizations split this further into MTTC (Mean Time to Contain) as a separate checkpoint before full remediation. | MTTR reflects how efficiently the SOC and its supporting teams move from "we know what's happening" to "it's handled." It's shaped by playbook quality, tooling, and how much cross-team coordination a given incident requires. |
Reading the Metrics Together
Read together, these three metrics describe a timeline: detect, acknowledge, respond. A weakness in any one stage drags down the overall picture even if the other two look strong. A SOC with an excellent MTTD but a poor MTTA has good detection engineering sitting behind an understaffed or overwhelmed triage queue; a SOC with fast MTTA but slow MTTR probably has a staffing or process problem downstream of triage, in containment or remediation rather than in noticing the alert.
Coverage Metrics
Time-based metrics describe how fast the SOC reacts to what it sees. Coverage metrics describe how much of the environment it's actually seeing and defending against in the first place, and they matter just as much, because a SOC can have an excellent MTTD on the alerts it does generate while remaining blind to entire categories of activity it never had visibility into.
Log Source Coverage
Log source coverage asks a simple question: of the critical assets and systems in the environment, how many are actually feeding relevant telemetry into the SIEM or detection stack. A server that never sends logs is a server the SOC cannot detect anything on, no matter how well-tuned the detection rules are elsewhere.
Tracking this conceptually, as a gap list of what's onboarded versus what's known to exist but isn't yet, is more useful than chasing a specific percentage figure, since the goal isn't a number on a dashboard but closing the actual visibility gaps that number represents.
Detection Coverage Against ATT&CK
MITRE ATT&CK gives a SOC a shared reference for mapping which adversary tactics and techniques it has built detections for, and which it hasn't. Plotting existing detection rules against the ATT&CK matrix produces a visual gap analysis: tactics with dense coverage, tactics with a single thin detection, and tactics with none at all.
This is the same mapping exercise covered from the hunting angle earlier in this learning path, applied here as an ongoing coverage metric rather than a one-time hunting exercise.
Used honestly, coverage metrics answer a different question than time-based metrics do. Time-based metrics tell a SOC how well it's handling what it catches. Coverage metrics tell it what it might not be catching at all, which is the harder and more uncomfortable of the two questions, and arguably the more important one to keep asking.
SOC Maturity Models
Individual metrics, even a well-chosen set of them, describe pieces of a SOC's performance. A maturity model steps back and assesses the program as a whole: not just how fast alerts get handled this quarter, but how repeatable, documented, and resilient the SOC's people, processes, technology, and services are as a system.
SOC-CMM
SOC-CMM (Security Operations Center Capability Maturity Model) is one well-known example of this kind of framework, built specifically for security operations rather than adapted from a generic IT maturity model. Frameworks in this family generally assess a SOC across a handful of dimensions:
- People: staffing, skills, training, retention
- Process: documented, repeatable procedures rather than tribal knowledge
- Technology: the tooling stack and how well it's actually configured and used, not just what's licensed
- Services: the specific capabilities the SOC delivers, such as monitoring, incident response, or threat hunting, and how consistently it delivers them
What Maturity Adds
What a maturity assessment adds that a metrics dashboard doesn't is context for those metrics. A SOC with a strong MTTR might still score low on maturity if that speed depends entirely on one senior analyst's institutional knowledge rather than a documented, repeatable process; that's a fragile strength, not a durable one, and it tends to show up the moment that analyst is out sick or leaves.
Maturity models are built to surface exactly that kind of gap between good current-state numbers and a program that will hold up over time.
Metrics Pitfalls
Any metric that people are evaluated against will eventually get optimized for, sometimes at the expense of the outcome the metric was supposed to represent. This isn't a hypothetical risk specific to security operations; it shows up anywhere a number becomes a target. A SOC's metrics program needs to be designed with that tendency in mind, not surprised by it after the fact.
Common Failure Modes
- Gaming metrics: an analyst under pressure to improve MTTR can close tickets faster by skipping proper investigation, producing a great-looking number and a worse-investigated case.
- Vanity metrics: raw alert volume handled looks impressive on a slide but says nothing about whether those dispositions were accurate; a team that closes a thousand alerts a day poorly is not outperforming one that closes two hundred correctly.
- Speed versus thoroughness: nearly every time-based metric creates pressure to move faster, and moving faster without a corresponding check on quality is how real incidents get missed or mis-triaged in the name of a better dashboard number.
Pairing Speed With Quality
The practical defense against all three of these is pairing every speed metric with a quality metric, rather than tracking time alone. Disposition accuracy (how often an analyst's initial call on an alert holds up on review) and reopened-ticket rate (how often a "resolved" case comes back because it wasn't actually handled) are two common pairings.
A team with a fast MTTR and a high reopened-ticket rate isn't actually fast; it's deferring the real work to a second pass, and the paired metric is what makes that visible instead of hidden behind an impressive-looking average.
None of this means time-based and coverage metrics aren't worth tracking. It means they're incomplete on their own, and a metrics program that leadership, the SOC, and auditors can all trust needs to show its work: not just how fast, but how accurately, and not just how much coverage, but how good that coverage actually is when tested.
Key Takeaways
- Metrics serve three distinct audiences: leadership (effectiveness and staffing justification), the SOC itself (a tuning feedback loop), and compliance/audit (evidence the program operates as claimed).
- MTTD, MTTA, and MTTR describe a timeline: detect, acknowledge, respond. Definitions vary between organizations, so compare against your own historical trend rather than an external benchmark.
- Coverage metrics, log source coverage and ATT&CK-mapped detection coverage, answer a different question than speed metrics: what might the SOC not be catching at all. Treat coverage mapping as gap analysis, not a score to game.
- SOC maturity models like SOC-CMM assess people, process, technology, and services as a system, surfacing fragile strengths (like a fast MTTR built on one analyst's knowledge) that a metrics dashboard alone would miss.
- Any metric people are evaluated against can be gamed. Pair time-based metrics with quality metrics, such as disposition accuracy or reopened-ticket rate, to catch speed gained at the expense of thoroughness.
Knowledge Check
Click an answer to reveal the explanation.
A SOC lead wants a number that helps justify a staffing request to executive leadership, while an auditor separately wants evidence the incident response process actually works. What does this illustrate about SOC metrics?
Why does this chapter caution against comparing your organization's MTTD, MTTA, or MTTR directly to a number published in a vendor report or another company's blog post?
A SOC's average MTTR improves significantly after leadership starts reviewing it monthly, but the reopened-ticket rate climbs at the same time. What does this pattern most likely indicate?