Why we did this

Almost every security standard tells businesses to patch by severity: fix anything scored 7 or above, usually within 14 days. It is written into Cyber Essentials, it is how most tools sort their queues, and for a business without a security team it produces a list far longer than anyone can work through.

So we asked a question the standards do not answer. If you can only act on some vulnerabilities, which ones actually matter? Not which score highest, but which turn out to be used against real businesses.

How we did it

The analysis covers 400,837 distinct public vulnerabilities, queried on 4 October 2026. The sources are public: the National Vulnerability Database, CISA's Known Exploited Vulnerabilities catalogue, GitHub Security Advisories, and the Exploit Prediction Scoring System. Exploitation is measured against CISA's catalogue, which records flaws confirmed as being used in the wild.

Three limits are worth stating plainly, because they shape everything below.

First, CISA's catalogue is a floor rather than a census. It records what one agency has confirmed and chosen to publish, mainly to drive remediation across US federal government. Plenty of real exploitation never reaches it, so every figure here understates rather than overstates.

Second, the severity scores we analysed are current ones, not the score a flaw carried on the day it was published. Scores get revised upwards once exploitation becomes known, so where we report that exploited flaws were scored low, the true figure at the time a defender had to decide is higher than what we can show.

Third, the record itself is incomplete. About a fifth of it carries no severity score at all, rising to more than two in five of this year's entries, and around 11% cannot be reliably matched to a product. Those gaps are findings in their own right, and they cap how far any product-level analysis can go.

The data was gathered using StackFlag, which is a Steelwise product. The underlying vulnerability data is public and third-party. We have published the caveats that weaken our findings alongside the findings themselves, including two defects in the collection that we found while doing this and had to correct before the analysis stood up.

Corrected on 4 October 2026. We later found a third defect: about 180,000 records from the National Vulnerability Database had silently failed to store, so the analysis first published here in July and August 2026 ran on an incomplete record of 248,176 entries. We re-ran everything on the complete data. Every figure on this page changed, two findings were withdrawn, and the advantage of the probability-based patching rule turned out to be much smaller than we first reported. Each filing carries its own correction note setting out what changed.

What we found

  • Sorting by severity is expensive and imprecise. Patching everything scored 7 or above means working through 198,716 items to catch 89% of exploited flaws, and 99.2% of that list was never exploited. A probability-based rule catches 74% from a list of 17,217, less than a tenth of the size.
  • "Just do the criticals" is the worst rule we tested. It catches 36% of what gets exploited, and still leaves 54,012 things to do. The probability-based rule catches twice as much from a third of the list.
  • 99% of critical-scored vulnerabilities are never exploited by anyone. Of 54,012 flaws scored 9.0 or higher, 626 show any exploitation evidence.
  • The mismatch runs both ways. About one in nine confirmed-exploited flaws was scored below 7, so a threshold filter never looks at them.
  • The typical exploited flaw is eight months old. Median time from publication to confirmed exploitation is 242 days, and 45.0% take more than a year. Watching only new vulnerabilities misses most of what matters.
  • Counting flaws per product is close to useless as a risk signal. The Linux kernel has 15,006 recorded vulnerabilities and 0.2% exploited. FortiOS has 284 and about one in twelve.

The filings, in reading order

These were published as four separate filings. They are best read in this order rather than by date.

  1. The patching rule that beats "just do the criticals". Start here: which selection rule actually catches exploited vulnerabilities, and what the alternative costs.
  2. 99% of critical vulnerabilities are never exploited by anyone. What a severity score measures, what it misses, and which parts of it carry real signal.
  3. The vulnerability that gets you is eight months old. The gap between publication and exploitation, and why the backlog matters more than the inbox.
  4. Counting vulnerabilities tells you almost nothing about risk. Why flaw counts mislead, and what to look at instead.

If you are deciding what to patch

The short version: meet whatever standard you are certifying against, then use exploitation probability to decide what order to work through it in. The first filing sets out the numbers behind that, and the third explains why the old items on your list deserve more attention than they usually get.

Questions about this research

If you are a journalist, researcher, or practitioner and want to discuss the method or the findings, email contact@steelwise.uk. If you are a business trying to work out what to patch, the contact page is the place to start.