CHAPTER 06 35 MIN READ INTERMEDIATE

Risk & Vulnerability Management

"Patch everything immediately" sounds like a security policy, but no organization with more than a handful of systems can actually run one. Every environment carries thousands of open findings at any given moment, and the systems that need to stay up cannot all go down for maintenance windows at the same time. Risk management exists because resources are finite and not every vulnerability deserves the same response. This chapter builds the mental model for making those tradeoffs deliberately: how risk is scored, how vulnerabilities move through a management lifecycle, and how real teams decide what gets fixed first, what gets a compensating control, and what gets formally accepted.

CVSS vuln lifecycle patch management

Risk as Likelihood Times Impact

The foundational model underneath every risk decision in security is deceptively simple: risk equals likelihood multiplied by impact. Neither number means much on its own. A finding with catastrophic impact but effectively zero likelihood is not a priority, and a finding with trivial impact but near-certain likelihood might still not be either. Risk only becomes meaningful when you hold both dimensions at once.

Quick glossary, keep these three straight:
  • Likelihood: the probability a threat actually materializes, whether that is a vulnerability getting exploited, a phishing email getting clicked, or a misconfigured bucket getting discovered by a scanner.
  • Impact: what happens if it does, data loss, downtime, regulatory fines, reputational damage.
  • Risk: the product of the two. High impact with near-zero likelihood is low risk; high likelihood with trivial impact is also low risk.

Why Impact Alone Misleads

This is where a lot of junior analysts get the model wrong. It is tempting to rank findings purely by impact, because impact is the scarier, more concrete number: "this could lead to full domain compromise" reads as more urgent than "this could leak a low-sensitivity config file."

But a critical-impact vulnerability sitting on an air-gapped system with no network path to it and no known exploit in the wild can carry lower actual risk than a medium-impact finding on an internet-facing login page that automated scanners are already probing. The first has near-zero likelihood. The second has likelihood approaching certainty. Prioritizing by impact alone routinely misallocates remediation effort toward things that look scary and away from things that are actually being attacked.

Why Likelihood Alone Misleads

The inverse mistake is just as common: treating likelihood as the only axis and chasing whatever is noisiest. A constant stream of low-impact alerts from a scanner can consume an entire team's bandwidth while a rare but catastrophic scenario, like a single unpatched internet-facing RCE with a public exploit, goes unaddressed because it only generates one line item.

Constant-but-minor and rare-but-severe both deserve attention, but they deserve different kinds of attention. Mistaking volume for severity is how teams end up busy without being effective.

What Feeds Each Axis

Likelihood is built from several inputs: is there a public exploit, is the vulnerable system exposed to the internet or only reachable internally, does exploitation require authentication or user interaction, and is there evidence the vulnerability class is actively being used in the wild.

Impact is built from asset criticality, data sensitivity, and blast radius, meaning how far an attacker could pivot from that one foothold. A prioritization model that only asks "how bad is this vulnerability" without also asking "how likely is someone to actually pull this off, against this specific asset" will consistently produce a list that feels rigorous but does not track real-world risk.

This likelihood-times-impact framing is the reasoning underneath every other tool in this chapter. CVSS base scores mostly measure impact and exploitability characteristics of the vulnerability itself, in isolation from your environment. The vulnerability management lifecycle exists to layer likelihood, in the form of exploit availability and exposure, back on top of that base score. Risk treatment decisions, accept, mitigate, transfer, or avoid, are really just different ways of responding once you know where a finding sits on both axes.

CVSS Scoring Explained

The Common Vulnerability Scoring System (CVSS) is the industry-standard method for expressing how severe a given vulnerability is, on a scale from 0.0 to 10.0. Almost every vulnerability scanner, CVE entry, and vendor advisory reports a CVSS base score, so understanding how that number is built is a prerequisite for using it correctly rather than just reading it as a single opaque digit. The base score is derived from a set of metrics that describe the intrinsic characteristics of the vulnerability itself, independent of any specific organization's environment.

Exploitability Metrics

Attack Vector describes how an attacker reaches the vulnerable component: Network, Adjacent, Local, or Physical. A vulnerability exploitable remotely over the internet scores higher on this metric than one that requires physical access to the device, because a network-reachable flaw has a vastly larger pool of potential attackers.

Attack Complexity captures whether exploitation is reliable and repeatable (Low) or depends on conditions outside the attacker's control, like winning a race condition or needing specific configuration to be present (High). Privileges Required asks whether the attacker needs no account, a low-privilege account, or administrative rights before the attack can even begin. User Interaction asks whether a victim has to do something, like click a link or open a file, for the exploit to succeed.

Those four metrics combine to describe exploitability: how easy is it, practically speaking, for an attacker to trigger this flaw.

Impact Metrics (CIA Triad)

The other half of the base score is impact, expressed across the classic CIA triad. Each of the three is rated None, Low, or High.

  • Confidentiality impact: how much unauthorized data disclosure results. A SQL injection that dumps an entire customer database scores High here.
  • Integrity impact: how much unauthorized data modification results.
  • Availability impact: how much the exploit degrades or denies access to the system. A denial-of-service flaw that only crashes a service scores High on availability but None on confidentiality and integrity.

The exploitability metrics and the impact metrics are combined through a defined formula to produce the base score.

Temporal and Environmental Adjustments

The base score is only the starting point. Temporal metrics adjust that score based on factors that change over time, most importantly whether working exploit code is publicly available and whether an official patch exists yet. A vulnerability with a public proof-of-concept exploit is more urgent today than the same vulnerability was the day it was disclosed with no exploit in the wild.

Environmental metrics go a step further and let an organization adjust the score based on their specific deployment: how critical is the affected asset to them, and what compensating controls, like network segmentation or WAF rules, already reduce the practical exploitability. A 9.8 base score CVE sitting behind three layers of network isolation with no direct exposure does not carry a 9.8 worth of real risk to that organization, which is exactly the likelihood-times-impact reasoning from the previous section applied through the CVSS environmental layer.

Vendors and scanners publish the base score by default because it is the one figure that stays constant and comparable across environments. Practitioners doing prioritization work should treat the base score as a starting input, not a final verdict, and layer temporal and environmental context on top before deciding what to patch first.

Score RangeSeverity Rating
0.1 – 3.9Low
4.0 – 6.9Medium
7.0 – 8.9High
9.0 – 10.0Critical

The Vulnerability Management Lifecycle

Vulnerability management is not a one-time project, it is a repeating cycle that runs continuously against an environment that never stops changing. New assets get deployed, new CVEs get disclosed, and configurations drift, so a scan result from three months ago tells you almost nothing about your current exposure. The lifecycle is commonly described in four stages, and the cycle feeds directly back into itself.

  1. Discover. Automated tools, whether network-based scanners, agent-based tools, or cloud configuration scanners, sweep the environment and identify known vulnerabilities, missing patches, and insecure configurations. Discovery quality depends heavily on coverage: a scan that misses a subnet, a set of unmanaged endpoints, or a shadow IT cloud account is not producing a complete picture, no matter how good the scanning engine is on the assets it does see. Asset inventory accuracy is a prerequisite for effective discovery, because you cannot scan what you do not know exists.
  2. Prioritize. Take the raw list of findings from discovery and rank them using the risk model from earlier in this chapter: CVSS base score as a starting point, adjusted by asset criticality (is this a domain controller or a test VM), exposure (internet-facing or internal-only), and real-world exploitability (is there a public exploit, is it in active use). A scanner might return thousands of findings from a single sweep of a mid-size environment, and nobody remediates thousands of findings in a week, so prioritization is what turns an unmanageable list into an actionable one.
  3. Remediate. This is where the finding actually gets addressed, and it is not limited to installing a patch. Remediation can mean applying a vendor patch, applying a compensating control like a firewall rule or WAF signature that blocks the exploitation path without touching the vulnerable code, or, for findings that fall below the organization's risk threshold, formally accepting the risk and documenting why. Which path is appropriate depends on the finding, the asset, and operational constraints, which is exactly what the next section covers in detail.
  4. Verify. After remediation, the asset gets rescanned to confirm the finding is actually resolved, not just marked closed in a ticketing system. This step catches the common failure where a patch was scheduled but never actually deployed, or where a config change reverted after the next system rebuild. Verification feeds straight back into discovery, closing the loop.

Organizations that treat vulnerability management as a linear project rather than this continuous loop tend to see their exposure quietly climb back up within a few months of any point-in-time remediation push. The next scan cycle picks up new findings introduced by whatever changed in the environment since the last pass, and the cycle repeats.

Patch Management Under Real Constraints

"Patch everything within 24 hours" is a policy that sounds rigorous in a slide deck and falls apart the first week it meets production infrastructure. Most organizations run some mix of change control processes, legacy systems that cannot tolerate certain updates, and uptime requirements that make "patch immediately" operationally impossible for a meaningful share of the environment. Understanding why matters more than memorizing the ideal SLA, because the real skill in patch management is triaging intelligently within those constraints, not pretending they do not exist.

Why Change Control Slows Things Down

Change control windows exist because an untested patch applied directly to production can cause an outage that costs more than the vulnerability it was meant to fix. Most mature organizations require patches to move through a staging or test environment first, then get scheduled into an approved maintenance window, which might only occur weekly or monthly for certain critical systems.

That review process is not laziness, it is a control against the very real risk of a patch breaking a dependent application, a driver, or an integration that nobody documented.

Legacy Systems Compound the Problem

A vendor may have stopped supporting a piece of software years ago, meaning no patch exists at all for a newly disclosed vulnerability in it, or the underlying OS may be so old that the vendor's current patch package silently assumes libraries or configurations that were never present. Applying a modern patch to that kind of system can break it outright.

If the affected system happens to run a manufacturing line, a medical device, or a piece of financial infrastructure that cannot simply be taken offline, "just patch it" is not a real option. These systems typically get handled through network isolation and compensating controls instead, which is a risk treatment decision covered in the next section.

Using Exploitability Data to Triage

Given those constraints, teams cannot treat every CVE as equally urgent, so they lean on exploitability data to decide what jumps the queue. CISA's Known Exploited Vulnerabilities (KEV) catalog is the most widely used input for this: it lists CVEs that are confirmed to be under active exploitation in the wild, not just theoretically exploitable.

A CVE landing on the KEV list is a strong signal to prioritize it ahead of other findings with similar CVSS scores that have no evidence of real-world use, because "actively exploited right now" is a much stronger likelihood signal than a base score alone can capture. Federal agencies under CISA's binding operational directives are required to remediate KEV-listed vulnerabilities on fixed timelines, and many private organizations adopt the same list as a practical prioritization filter even without that mandate.

Warning: A high CVSS score and active exploitation are not the same signal. A 9.8 CVE with no public exploit and no KEV listing can reasonably wait behind a 7.5 CVE that is actively being used in ransomware campaigns. Prioritization models that only sort by base score miss this distinction and can leave the actually dangerous findings unpatched while resources go toward theoretical worst-case scenarios.

Risk Treatment: Accept, Mitigate, Transfer, Avoid

Once a finding has been discovered and prioritized, it does not automatically follow that the answer is "patch it." Formal risk management defines four treatment strategies, and choosing among them deliberately, rather than defaulting to whichever one is easiest, is what separates a mature program from one that just closes tickets.

StrategyWhat It MeansConcrete Example
AcceptFormally acknowledge the risk and decide, with sign-off, not to act on it right nowA medium-severity finding on an internal test server with no sensitive data, no path to production, scheduled for decommission in two months
MitigateReduce the risk without fully eliminating the underlying vulnerability, usually via a compensating controlAn unsupported legacy app placed behind a WAF rule, with network access restricted and enhanced monitoring added
TransferShift the financial consequence to a third party, most commonly cyber insuranceA policy that partially covers breach notification, legal fees, or business interruption costs if the risk materializes
AvoidEliminate the risk entirely by removing the thing that creates itDecommissioning an unsupported legacy system outright and migrating its function to a supported platform

Accept

Accept means the organization formally acknowledges the risk and decides, with sign-off from someone accountable for that decision, not to act on it right now. This is appropriate for low-likelihood or low-impact findings where the cost of remediation outweighs the exposure.

A risk owner documents the acceptance, including the reasoning and an expiration date for review, rather than letting the finding sit as an unaddressed red item in a dashboard indefinitely.

Mitigate

Mitigate means reducing the risk without fully eliminating the underlying vulnerability, usually through a compensating control. If a legacy application cannot be patched because the vendor stopped supporting it, a team might place it behind a web application firewall with a rule that blocks the specific exploitation pattern, restrict network access to only the handful of systems that legitimately need to reach it, and add enhanced monitoring for anomalous activity against that host.

The vulnerability still technically exists, but the practical likelihood of successful exploitation drops substantially.

Transfer

Transfer means shifting the financial consequence of the risk to a third party, most commonly through cyber insurance. The vulnerability or exposure itself is not fixed or reduced, but if it is exploited, the resulting costs, such as breach notification, legal fees, or business interruption losses, are partially covered by the policy.

Transfer is common for risks that are expensive to fully remediate or inherent to doing business in a particular way. It is typically used alongside other treatments rather than as a sole strategy, since insurers increasingly require baseline security controls before they will underwrite a policy at all.

Avoid

Avoid means eliminating the risk entirely by removing the thing that creates it. If a legacy system is running an unsupported OS with no patch path and no reasonable compensating control that reduces risk to an acceptable level, the organization may decide to decommission it outright and migrate its function to a supported platform.

This is often the most effective treatment and also the most expensive and disruptive, which is why it tends to be reserved for findings where mitigation and acceptance both fall short of an acceptable risk level.

Vulnerability Scanning vs Penetration Testing

Vulnerability scanning and penetration testing get lumped together as "security testing" in casual conversation, but they answer different questions and neither substitutes for the other. Confusing the two leads to gaps: organizations that only scan miss the exploitation chains a human attacker would find, and organizations that only pentest miss the continuous drift that scanning is built to catch.

AspectVulnerability ScanningPenetration Testing
NatureBroad, shallow, continuousNarrow, deep, point-in-time
MethodAutomated tool checks systems against a CVE/misconfiguration databaseHuman tester actively exploits and chains vulnerabilities like a real adversary
CadenceScheduled, weekly or even daily for critical assetsOnce or twice a year, against a defined scope and timeframe
ScaleCovers thousands of hosts with no added headcountLimited to what fits the engagement scope
Key limitationDoes not chain findings together or attempt actual exploitationNot running often or broadly enough to catch next Tuesday's CVE on an out-of-scope system

What Scanning Misses

Vulnerability scanning produces a list of findings, typically scored by something like CVSS, but it generally does not attempt actual exploitation. It can report a vulnerability exists without confirming that exploiting it is actually feasible in that specific environment.

What Pentesting Adds

A human tester works toward gaining an initial foothold, escalating privileges, pivoting to other systems, and demonstrating actual business impact rather than just theoretical risk. A pentest might take a single low-severity misconfiguration that a scanner would rate as minor and combine it with two other minor findings to achieve full domain compromise, a chain a scanner cannot construct on its own because it requires creative, adversarial human reasoning.

Why Programs Need Both

The tradeoff is coverage versus depth. Scanning cannot substitute for a pentest either, because it will not tell you that three separate low-severity findings combine into a critical attack path, and it will not test how well your detection and response actually holds up against a skilled, motivated human working to evade it.

In practice, mature programs run both, using them for what each is good at: scanning as the continuous baseline that drives the vulnerability management lifecycle described earlier, and periodic penetration testing, often required for compliance frameworks like PCI DSS, as a deeper validation that surfaces the exploitation chains and business-impact scenarios that scanning alone cannot see.

Tip: If you are asked to explain the difference in an interview, lead with the tradeoff: scanning is automated, broad, and continuous; pentesting is manual, narrow, and point-in-time. Then give one concrete example of a finding a scanner would report versus a chain a pentester would build, that framing demonstrates you understand the practical difference, not just the definitions.

Key Takeaways

  • Risk equals likelihood times impact. Prioritizing by impact alone or likelihood alone both produce distorted, ineffective remediation queues.
  • CVSS base scores combine exploitability metrics (Attack Vector, Attack Complexity, Privileges Required, User Interaction) with CIA impact metrics to produce a 0-10 score; temporal and environmental metrics adjust that base score for real-world context.
  • The vulnerability management lifecycle repeats continuously: discover through scanning, prioritize using risk and exploitability data, remediate through patching or compensating controls, and verify through rescanning.
  • Change control windows, legacy system fragility, and uptime requirements make "patch everything immediately" operationally impossible, which is why exploitability data like CISA's KEV catalog is used to triage what jumps the queue.
  • The four risk treatment strategies are accept, mitigate, transfer, and avoid; choosing among them deliberately, with documented ownership, is what separates mature risk management from just closing tickets.
  • Vulnerability scanning is broad, shallow, and continuous; penetration testing is narrow, deep, and point-in-time. Mature programs run both because each catches what the other misses.

Knowledge Check

Click an answer to reveal the explanation.

A scanner reports two findings: a CVSS 9.8 critical vulnerability on an isolated lab server with no network path to production, and a CVSS 7.2 high vulnerability on an internet-facing login page that is listed in CISA's KEV catalog as actively exploited. Which should be remediated first?

Risk is likelihood times impact, not impact alone. The isolated lab server has near-zero likelihood of exploitation due to lack of network exposure, while the internet-facing, actively-exploited finding has high likelihood despite a lower base score. KEV listing and internet exposure are exactly the kind of real-world context that should override a raw CVSS comparison.

A team wants to know whether a set of three separately low-severity misconfigurations in their environment could be chained together by an attacker to reach a critical system. Which activity is best suited to answer that question?

Vulnerability scanners evaluate findings individually and do not construct multi-step exploitation chains the way a human attacker would. A penetration test is designed specifically to test whether findings can be combined to achieve meaningful impact, which is exactly the depth-over-breadth strength that distinguishes pentesting from automated scanning.

A legacy manufacturing control system cannot be patched because the vendor no longer supports it, and taking it offline would halt production. The security team places it on an isolated VLAN, restricts access to two authorized engineering workstations, and adds monitoring for anomalous traffic to it. This is an example of which risk treatment strategy?

The underlying vulnerability is not fixed, and the system is not being decommissioned or accepted without action. Compensating controls, network isolation, access restriction, and monitoring, are being layered on to reduce the practical likelihood of exploitation, which is the definition of mitigation. Accept would mean documenting and doing nothing; avoid would mean decommissioning the system entirely.