Risk & Vulnerability Management
"Patch everything immediately" sounds like a security policy, but no organization with more than a handful of systems can actually run one. Every environment carries thousands of open findings at any given moment, and the systems that need to stay up cannot all go down for maintenance windows at the same time. Risk management exists because resources are finite and not every vulnerability deserves the same response. This chapter builds the mental model for making those tradeoffs deliberately: how risk is scored, how vulnerabilities move through a management lifecycle, and how real teams decide what gets fixed first, what gets a compensating control, and what gets formally accepted.
Risk as Likelihood Times Impact
The foundational model underneath every risk decision in security is deceptively simple: risk equals likelihood multiplied by impact. Neither number means much on its own. A finding with catastrophic impact but effectively zero likelihood is not a priority, and a finding with trivial impact but near-certain likelihood might still not be either. Risk only becomes meaningful when you hold both dimensions at once.
- Likelihood: the probability a threat actually materializes, whether that is a vulnerability getting exploited, a phishing email getting clicked, or a misconfigured bucket getting discovered by a scanner.
- Impact: what happens if it does, data loss, downtime, regulatory fines, reputational damage.
- Risk: the product of the two. High impact with near-zero likelihood is low risk; high likelihood with trivial impact is also low risk.
Why Impact Alone Misleads
This is where a lot of junior analysts get the model wrong. It is tempting to rank findings purely by impact, because impact is the scarier, more concrete number: "this could lead to full domain compromise" reads as more urgent than "this could leak a low-sensitivity config file."
But a critical-impact vulnerability sitting on an air-gapped system with no network path to it and no known exploit in the wild can carry lower actual risk than a medium-impact finding on an internet-facing login page that automated scanners are already probing. The first has near-zero likelihood. The second has likelihood approaching certainty. Prioritizing by impact alone routinely misallocates remediation effort toward things that look scary and away from things that are actually being attacked.
Why Likelihood Alone Misleads
The inverse mistake is just as common: treating likelihood as the only axis and chasing whatever is noisiest. A constant stream of low-impact alerts from a scanner can consume an entire team's bandwidth while a rare but catastrophic scenario, like a single unpatched internet-facing RCE with a public exploit, goes unaddressed because it only generates one line item.
Constant-but-minor and rare-but-severe both deserve attention, but they deserve different kinds of attention. Mistaking volume for severity is how teams end up busy without being effective.
What Feeds Each Axis
Likelihood is built from several inputs: is there a public exploit, is the vulnerable system exposed to the internet or only reachable internally, does exploitation require authentication or user interaction, and is there evidence the vulnerability class is actively being used in the wild.
Impact is built from asset criticality, data sensitivity, and blast radius, meaning how far an attacker could pivot from that one foothold. A prioritization model that only asks "how bad is this vulnerability" without also asking "how likely is someone to actually pull this off, against this specific asset" will consistently produce a list that feels rigorous but does not track real-world risk.
This likelihood-times-impact framing is the reasoning underneath every other tool in this chapter. CVSS base scores mostly measure impact and exploitability characteristics of the vulnerability itself, in isolation from your environment. The vulnerability management lifecycle exists to layer likelihood, in the form of exploit availability and exposure, back on top of that base score. Risk treatment decisions, accept, mitigate, transfer, or avoid, are really just different ways of responding once you know where a finding sits on both axes.
CVSS Scoring Explained
The Common Vulnerability Scoring System (CVSS) is the industry-standard method for expressing how severe a given vulnerability is, on a scale from 0.0 to 10.0. Almost every vulnerability scanner, CVE entry, and vendor advisory reports a CVSS base score, so understanding how that number is built is a prerequisite for using it correctly rather than just reading it as a single opaque digit. The base score is derived from a set of metrics that describe the intrinsic characteristics of the vulnerability itself, independent of any specific organization's environment.
Exploitability Metrics
Attack Vector describes how an attacker reaches the vulnerable component: Network, Adjacent, Local, or Physical. A vulnerability exploitable remotely over the internet scores higher on this metric than one that requires physical access to the device, because a network-reachable flaw has a vastly larger pool of potential attackers.
Attack Complexity captures whether exploitation is reliable and repeatable (Low) or depends on conditions outside the attacker's control, like winning a race condition or needing specific configuration to be present (High). Privileges Required asks whether the attacker needs no account, a low-privilege account, or administrative rights before the attack can even begin. User Interaction asks whether a victim has to do something, like click a link or open a file, for the exploit to succeed.
Those four metrics combine to describe exploitability: how easy is it, practically speaking, for an attacker to trigger this flaw.
Impact Metrics (CIA Triad)
The other half of the base score is impact, expressed across the classic CIA triad. Each of the three is rated None, Low, or High.
- Confidentiality impact: how much unauthorized data disclosure results. A SQL injection that dumps an entire customer database scores High here.
- Integrity impact: how much unauthorized data modification results.
- Availability impact: how much the exploit degrades or denies access to the system. A denial-of-service flaw that only crashes a service scores High on availability but None on confidentiality and integrity.
The exploitability metrics and the impact metrics are combined through a defined formula to produce the base score.
Temporal and Environmental Adjustments
The base score is only the starting point. Temporal metrics adjust that score based on factors that change over time, most importantly whether working exploit code is publicly available and whether an official patch exists yet. A vulnerability with a public proof-of-concept exploit is more urgent today than the same vulnerability was the day it was disclosed with no exploit in the wild.
Environmental metrics go a step further and let an organization adjust the score based on their specific deployment: how critical is the affected asset to them, and what compensating controls, like network segmentation or WAF rules, already reduce the practical exploitability. A 9.8 base score CVE sitting behind three layers of network isolation with no direct exposure does not carry a 9.8 worth of real risk to that organization, which is exactly the likelihood-times-impact reasoning from the previous section applied through the CVSS environmental layer.
Vendors and scanners publish the base score by default because it is the one figure that stays constant and comparable across environments. Practitioners doing prioritization work should treat the base score as a starting input, not a final verdict, and layer temporal and environmental context on top before deciding what to patch first.
| Score Range | Severity Rating |
|---|---|
| 0.1 – 3.9 | Low |
| 4.0 – 6.9 | Medium |
| 7.0 – 8.9 | High |
| 9.0 – 10.0 | Critical |
The Vulnerability Management Lifecycle
Vulnerability management is not a one-time project, it is a repeating cycle that runs continuously against an environment that never stops changing. New assets get deployed, new CVEs get disclosed, and configurations drift, so a scan result from three months ago tells you almost nothing about your current exposure. The lifecycle is commonly described in four stages, and the cycle feeds directly back into itself.
- Discover. Automated tools, whether network-based scanners, agent-based tools, or cloud configuration scanners, sweep the environment and identify known vulnerabilities, missing patches, and insecure configurations. Discovery quality depends heavily on coverage: a scan that misses a subnet, a set of unmanaged endpoints, or a shadow IT cloud account is not producing a complete picture, no matter how good the scanning engine is on the assets it does see. Asset inventory accuracy is a prerequisite for effective discovery, because you cannot scan what you do not know exists.
- Prioritize. Take the raw list of findings from discovery and rank them using the risk model from earlier in this chapter: CVSS base score as a starting point, adjusted by asset criticality (is this a domain controller or a test VM), exposure (internet-facing or internal-only), and real-world exploitability (is there a public exploit, is it in active use). A scanner might return thousands of findings from a single sweep of a mid-size environment, and nobody remediates thousands of findings in a week, so prioritization is what turns an unmanageable list into an actionable one.
- Remediate. This is where the finding actually gets addressed, and it is not limited to installing a patch. Remediation can mean applying a vendor patch, applying a compensating control like a firewall rule or WAF signature that blocks the exploitation path without touching the vulnerable code, or, for findings that fall below the organization's risk threshold, formally accepting the risk and documenting why. Which path is appropriate depends on the finding, the asset, and operational constraints, which is exactly what the next section covers in detail.
- Verify. After remediation, the asset gets rescanned to confirm the finding is actually resolved, not just marked closed in a ticketing system. This step catches the common failure where a patch was scheduled but never actually deployed, or where a config change reverted after the next system rebuild. Verification feeds straight back into discovery, closing the loop.
Organizations that treat vulnerability management as a linear project rather than this continuous loop tend to see their exposure quietly climb back up within a few months of any point-in-time remediation push. The next scan cycle picks up new findings introduced by whatever changed in the environment since the last pass, and the cycle repeats.
Patch Management Under Real Constraints
"Patch everything within 24 hours" is a policy that sounds rigorous in a slide deck and falls apart the first week it meets production infrastructure. Most organizations run some mix of change control processes, legacy systems that cannot tolerate certain updates, and uptime requirements that make "patch immediately" operationally impossible for a meaningful share of the environment. Understanding why matters more than memorizing the ideal SLA, because the real skill in patch management is triaging intelligently within those constraints, not pretending they do not exist.
Why Change Control Slows Things Down
Change control windows exist because an untested patch applied directly to production can cause an outage that costs more than the vulnerability it was meant to fix. Most mature organizations require patches to move through a staging or test environment first, then get scheduled into an approved maintenance window, which might only occur weekly or monthly for certain critical systems.
That review process is not laziness, it is a control against the very real risk of a patch breaking a dependent application, a driver, or an integration that nobody documented.
Legacy Systems Compound the Problem
A vendor may have stopped supporting a piece of software years ago, meaning no patch exists at all for a newly disclosed vulnerability in it, or the underlying OS may be so old that the vendor's current patch package silently assumes libraries or configurations that were never present. Applying a modern patch to that kind of system can break it outright.
If the affected system happens to run a manufacturing line, a medical device, or a piece of financial infrastructure that cannot simply be taken offline, "just patch it" is not a real option. These systems typically get handled through network isolation and compensating controls instead, which is a risk treatment decision covered in the next section.
Using Exploitability Data to Triage
Given those constraints, teams cannot treat every CVE as equally urgent, so they lean on exploitability data to decide what jumps the queue. CISA's Known Exploited Vulnerabilities (KEV) catalog is the most widely used input for this: it lists CVEs that are confirmed to be under active exploitation in the wild, not just theoretically exploitable.
A CVE landing on the KEV list is a strong signal to prioritize it ahead of other findings with similar CVSS scores that have no evidence of real-world use, because "actively exploited right now" is a much stronger likelihood signal than a base score alone can capture. Federal agencies under CISA's binding operational directives are required to remediate KEV-listed vulnerabilities on fixed timelines, and many private organizations adopt the same list as a practical prioritization filter even without that mandate.
Risk Treatment: Accept, Mitigate, Transfer, Avoid
Once a finding has been discovered and prioritized, it does not automatically follow that the answer is "patch it." Formal risk management defines four treatment strategies, and choosing among them deliberately, rather than defaulting to whichever one is easiest, is what separates a mature program from one that just closes tickets.
| Strategy | What It Means | Concrete Example |
|---|---|---|
| Accept | Formally acknowledge the risk and decide, with sign-off, not to act on it right now | A medium-severity finding on an internal test server with no sensitive data, no path to production, scheduled for decommission in two months |
| Mitigate | Reduce the risk without fully eliminating the underlying vulnerability, usually via a compensating control | An unsupported legacy app placed behind a WAF rule, with network access restricted and enhanced monitoring added |
| Transfer | Shift the financial consequence to a third party, most commonly cyber insurance | A policy that partially covers breach notification, legal fees, or business interruption costs if the risk materializes |
| Avoid | Eliminate the risk entirely by removing the thing that creates it | Decommissioning an unsupported legacy system outright and migrating its function to a supported platform |
Accept
Accept means the organization formally acknowledges the risk and decides, with sign-off from someone accountable for that decision, not to act on it right now. This is appropriate for low-likelihood or low-impact findings where the cost of remediation outweighs the exposure.
A risk owner documents the acceptance, including the reasoning and an expiration date for review, rather than letting the finding sit as an unaddressed red item in a dashboard indefinitely.
Mitigate
Mitigate means reducing the risk without fully eliminating the underlying vulnerability, usually through a compensating control. If a legacy application cannot be patched because the vendor stopped supporting it, a team might place it behind a web application firewall with a rule that blocks the specific exploitation pattern, restrict network access to only the handful of systems that legitimately need to reach it, and add enhanced monitoring for anomalous activity against that host.
The vulnerability still technically exists, but the practical likelihood of successful exploitation drops substantially.
Transfer
Transfer means shifting the financial consequence of the risk to a third party, most commonly through cyber insurance. The vulnerability or exposure itself is not fixed or reduced, but if it is exploited, the resulting costs, such as breach notification, legal fees, or business interruption losses, are partially covered by the policy.
Transfer is common for risks that are expensive to fully remediate or inherent to doing business in a particular way. It is typically used alongside other treatments rather than as a sole strategy, since insurers increasingly require baseline security controls before they will underwrite a policy at all.
Avoid
Avoid means eliminating the risk entirely by removing the thing that creates it. If a legacy system is running an unsupported OS with no patch path and no reasonable compensating control that reduces risk to an acceptable level, the organization may decide to decommission it outright and migrate its function to a supported platform.
This is often the most effective treatment and also the most expensive and disruptive, which is why it tends to be reserved for findings where mitigation and acceptance both fall short of an acceptable risk level.
Vulnerability Scanning vs Penetration Testing
Vulnerability scanning and penetration testing get lumped together as "security testing" in casual conversation, but they answer different questions and neither substitutes for the other. Confusing the two leads to gaps: organizations that only scan miss the exploitation chains a human attacker would find, and organizations that only pentest miss the continuous drift that scanning is built to catch.
| Aspect | Vulnerability Scanning | Penetration Testing |
|---|---|---|
| Nature | Broad, shallow, continuous | Narrow, deep, point-in-time |
| Method | Automated tool checks systems against a CVE/misconfiguration database | Human tester actively exploits and chains vulnerabilities like a real adversary |
| Cadence | Scheduled, weekly or even daily for critical assets | Once or twice a year, against a defined scope and timeframe |
| Scale | Covers thousands of hosts with no added headcount | Limited to what fits the engagement scope |
| Key limitation | Does not chain findings together or attempt actual exploitation | Not running often or broadly enough to catch next Tuesday's CVE on an out-of-scope system |
What Scanning Misses
Vulnerability scanning produces a list of findings, typically scored by something like CVSS, but it generally does not attempt actual exploitation. It can report a vulnerability exists without confirming that exploiting it is actually feasible in that specific environment.
What Pentesting Adds
A human tester works toward gaining an initial foothold, escalating privileges, pivoting to other systems, and demonstrating actual business impact rather than just theoretical risk. A pentest might take a single low-severity misconfiguration that a scanner would rate as minor and combine it with two other minor findings to achieve full domain compromise, a chain a scanner cannot construct on its own because it requires creative, adversarial human reasoning.
Why Programs Need Both
The tradeoff is coverage versus depth. Scanning cannot substitute for a pentest either, because it will not tell you that three separate low-severity findings combine into a critical attack path, and it will not test how well your detection and response actually holds up against a skilled, motivated human working to evade it.
In practice, mature programs run both, using them for what each is good at: scanning as the continuous baseline that drives the vulnerability management lifecycle described earlier, and periodic penetration testing, often required for compliance frameworks like PCI DSS, as a deeper validation that surfaces the exploitation chains and business-impact scenarios that scanning alone cannot see.
Key Takeaways
- Risk equals likelihood times impact. Prioritizing by impact alone or likelihood alone both produce distorted, ineffective remediation queues.
- CVSS base scores combine exploitability metrics (Attack Vector, Attack Complexity, Privileges Required, User Interaction) with CIA impact metrics to produce a 0-10 score; temporal and environmental metrics adjust that base score for real-world context.
- The vulnerability management lifecycle repeats continuously: discover through scanning, prioritize using risk and exploitability data, remediate through patching or compensating controls, and verify through rescanning.
- Change control windows, legacy system fragility, and uptime requirements make "patch everything immediately" operationally impossible, which is why exploitability data like CISA's KEV catalog is used to triage what jumps the queue.
- The four risk treatment strategies are accept, mitigate, transfer, and avoid; choosing among them deliberately, with documented ownership, is what separates mature risk management from just closing tickets.
- Vulnerability scanning is broad, shallow, and continuous; penetration testing is narrow, deep, and point-in-time. Mature programs run both because each catches what the other misses.
Knowledge Check
Click an answer to reveal the explanation.
A scanner reports two findings: a CVSS 9.8 critical vulnerability on an isolated lab server with no network path to production, and a CVSS 7.2 high vulnerability on an internet-facing login page that is listed in CISA's KEV catalog as actively exploited. Which should be remediated first?
A team wants to know whether a set of three separately low-severity misconfigurations in their environment could be chained together by an attacker to reach a critical system. Which activity is best suited to answer that question?
A legacy manufacturing control system cannot be patched because the vendor no longer supports it, and taking it offline would halt production. The security team places it on an isolated VLAN, restricts access to two authorized engineering workstations, and adds monitoring for anomalous traffic to it. This is an example of which risk treatment strategy?