Nobody Ever Checked Whether You Fixed It
Key Intel / TL;DR
  • The industry measures assessment maturity by how often you test, which says nothing about whether anything got fixed.
  • Most organizations cannot answer how many findings from their last test are still open, because the tracking ends at handoff.
  • Findings decay through three mechanisms: severity gets argued down, ownership gets diffused, and the fix never gets verified.
  • Retest the fix, not the finding, and keep findings in the system your engineers already work from.
  • Report closure rate and time to close alongside the finding count, because a count with no closure figure is only half a measurement.

Here is a question worth asking your team this week. Of the findings in your last penetration test report, how many are closed?

Not how many were assigned, and not how many have a ticket open somewhere. Closed, meaning somebody made a change and somebody else confirmed the change worked.

In most organizations the honest answer is that nobody knows, and the reason is structural rather than negligent. The engagement was scoped, sold, executed, and reported, and every part of that process had an owner. What happens after the report had none.

What the Industry Measures Instead

Assessment maturity gets discussed almost entirely in terms of frequency and coverage. Annually or quarterly, internal and external, whether you do application testing, whether you do social engineering, and whether the scope includes the cloud environment.

Those are reasonable questions and they all describe the input. None of them describes whether the organization is any less exposed than it was before the money was spent.

We wrote recently about the part of the environment somebody quietly excluded before the work started, which is the front end of this same problem. That piece is about what the test was allowed to look at. This one is about what happened to what it found. Between them sits an engagement that is well governed at both edges and unmeasured through the middle.

The Three Ways a Finding Dies

Findings rarely get ignored outright. They decay, through mechanisms that each look reasonable in isolation.

Severity gets negotiated

A finding arrives rated high. Somebody who owns the affected system reads it and explains why the rating is too aggressive in context, because the system is not internet-facing, or the exploit requires an authenticated session, or there is a compensating control the tester did not know about.

Frequently that person is right. Testers work with incomplete context and severity ratings are estimates. The problem is that the conversation only ever runs one direction. I have sat in a great many finding reviews and I have never once watched a severity get argued upward.

Each individual downgrade is defensible. The aggregate is a report that started with eleven highs and ends the quarter with two, and the eleven is the number that went to the board.

Ownership diffuses

A finding gets assigned to a team rather than a person. The team has a backlog that was full before the report arrived, and the finding enters it at whatever priority the severity implies after the negotiation above.

Then the sprint runs, the finding does not make it, and it rolls. Nobody decided to defer it. It simply lost, repeatedly, to work that had a customer attached. Six months later it is in the backlog with a creation date that nobody looks at.

The fix is never verified

This is the one that matters most and gets the least attention. Somebody makes a change, marks the ticket resolved, and the finding is counted as closed. No one goes back and attempts the original attack path.

A meaningful share of those fixes do not work. The configuration was applied to the wrong environment, the patch was installed but the service was never restarted, the input validation handles the example in the report and not the underlying class, or the change was reverted three weeks later by a deployment that predated it. Our guidance on documenting a deferred patch so it reads as a decision covers the honest version of not fixing something. An unverified fix is worse than a documented deferral, because the deferral is at least accurate.

What Actually Closes the Loop

None of this requires a bigger testing budget, and most of it costs considerably less than the test itself did. Four of the five items below are process changes somebody can make this quarter without asking anybody for money.

Buy the retest in the original contract

Negotiate remediation verification into the engagement rather than treating it as a separate purchase, because a separate purchase requires somebody to justify spending again on work already paid for. Most firms will include a retest window of thirty to ninety days at little or no additional cost if you ask during procurement, and almost none will offer it.

Then use it. The retest window expiring unused is the most common outcome, and it expires because the fixes were not finished in time, which is itself the finding.

Retest the fix rather than the finding

Ask the tester to attempt the original attack path again, not to confirm that a setting now reads the way it should. Those are different tests and they produce different results.

A configuration review verifies that somebody changed something. An attempt verifies that the change accomplished what it was for. The gap between those two is where most false closures live.

Keep findings where the engineers already are

A finding tracked in a spreadsheet maintained by the security team is a finding tracked outside the system where work actually gets prioritized. Push them into the same backlog as everything else, tagged so you can query them, and accept that they will compete with feature work, because they were always competing with feature work and the spreadsheet only hid it.

What you gain is a real creation date, a real assignee, and the ability to answer the question at the top of this article in about ten seconds.

Record the severity argument

When a rating gets negotiated down, write down who asked, what the reasoning was, and who agreed. Not to assign blame, and it does two useful things. It makes the pattern visible if one team’s findings are consistently reduced, and it gives you a record if the downgraded finding is later the one that matters.

This takes two sentences per finding and it is the single cheapest control in this entire article. It is also the one most likely to be skipped, because writing it down feels like an accusation when it is only a record.

Report closure, not count

Whatever you currently report upward about assessments, add two figures. The percentage of findings from the previous engagement that are verifiably closed, and the median time from report to verified closure.

Those numbers will be worse than you expect the first time you produce them. That is the point, and the first honest measurement is the most valuable one you will ever take, because everything after it is a trend.

Why This Is Worth the Attention

The uncomfortable version of the argument is that an organization testing annually with a thirty percent verified closure rate is in worse shape than one testing every two years and closing everything. The first spends more, reports more activity, and carries more unresolved exposure, and the reporting it produces looks healthier.

That is the failure this creates. A finding count is a measure of how well the test worked. A closure rate is a measure of how well the organization works, and only one of those is the thing you were trying to buy.

If you want help building the tracking side of this, or an independent retest against findings you believe are closed, contact Grab The Axe. You can also take our free Human Attack Surface Score to see where the people-shaped gaps sit.

Chris Armour is Director of Information Security at Grab The Axe.

Distribute Intel
Chris Armour
Director of Information Security
Chris Armour
The Breaker & Builder.

Operating on the philosophy that 'you can't build a secure system if you don't know how to break it,' Chris leads our engineering division. A top 1% National Cyber League competitor, he hardens our digital infrastructure against the very exploits he has mastered.

View Author Page →