PXL Security LTD, Sofia, Bulgaria Offensive security since 2014[email protected]
MethodologyReportingDefence

Remediation decay: what we found when we retested last year's findings

By PXL Security29 September 202619 min read

Remediation is treated as a one-way door: a finding is reported, a ticket is closed, and the issue is assumed to stay fixed. It does not. When we returned to a client a year later and re-tested every prior finding, roughly half were still exploitable — some never fixed, some fixed and since undone. The retest is the only honest measure of whether a security programme is working, and most organisations have never run one.

The assumption nobody tests

The lifecycle most organisations actually run looks like this. A test is commissioned. A report arrives with findings, each with a severity and a recommendation. Someone converts them into tickets, which are triaged, assigned, argued about, deprioritised, re-prioritised, and eventually moved to a closed state. A spreadsheet turns amber, then green. The matter is considered settled until the next annual test.

Notice what is missing. At no point does anyone independently attempt the original attack again. What marks a finding as resolved is a state change in a ticketing system, made by the team that owns the code, on their own assessment of their own change. That is not verification; it is a self-assessment recorded in a tool built to track work, not to establish security facts.

Ticket closure and verified remediation are different events, and they diverge for entirely ordinary, non-negligent reasons:

  • A developer fixes the behaviour they understood from the report summary, which is not always the behaviour the tester exploited.
  • A fix is merged and the ticket closed on merge — before the change reaches the tested environment, or all of them.
  • The ticket is closed as won't fix, accepted risk or mitigated by compensating control, and months later that nuance has been flattened into "closed" in the reporting leadership sees.
  • A fix addresses the single instance named in the report, and nobody went looking for the other seven.

Almost nothing in the usual process distinguishes these outcomes from a genuine, durable fix. The tracker cannot tell the difference; it only records what a human typed into it. And because the next test is usually a fresh engagement with a fresh scope, last year's findings are rarely re-attempted deliberately. If one surfaces again, it tends to be logged as a new finding rather than as evidence that remediation failed — a quieter and much more expensive conclusion.

What a repeat engagement showed

On returning to a client for a second annual engagement, we did what we now do by default: before testing anything new, we re-ran every finding from the prior year's report. Not a document review, not a conversation about what had been done — an attempt, by a tester, to reproduce each original issue against the current system, using the original evidence as the specification.

There were roughly 23 prior findings. Each was re-tested and given an explicit disposition. The approximate shape of the result:

Prior-year findings re-tested      ~23

Still present                      ~half
Genuinely fixed                    ~a third
Not reproducible                   remainder
Environment-blocked                remainder

Roughly half were still present — still reachable, still exploitable, mostly by the same steps as the year before. About a third were genuinely fixed: the attack was attempted, it failed, and it failed for a reason we could attribute to a specific change rather than to luck or inaccessibility. The remainder either could not be reproduced, for reasons short of evidence of a fix, or could not be reached because the environment differed materially from the one originally tested.

The crucial discipline: every still-present finding was re-demonstrated with fresh evidence. We carried nothing forward on paper. A finding appearing in two consecutive reports because someone copied it across is an administrative artefact; one appearing twice with two independent sets of evidence, a year apart, is a fact about the system. The first is noise; the second is a conversation leadership can act on.

What made the exercise uncomfortable was not the still-present half. It was that the client's own tracking showed nearly all of those findings as closed. No deception was involved and nobody had behaved badly. The process simply had no step capable of catching the gap between "we did something" and "this no longer works".

Note. This is one engagement, one client, approximately 23 findings. The sample is small, the proportions are approximate, and none of it should be read as an industry rate for remediation failure. It illustrates a mechanism — findings decay, and nothing in a typical process notices — not a statistic. The useful number is not ours; it is the one you get from re-testing your own last report.

Four ways a finding comes back

"Still present" is one disposition covering four quite different failures, worth separating because they have different owners and different remedies.

1. Never actually fixed. The ticket was closed on intent, not on change. Someone read the finding, agreed with it, planned the work — and the work was then descoped, deferred into a backlog later bulk-closed, or folded into a refactor that never shipped. Sometimes a mitigation was discussed in a comment thread and the discussion became the resolution. The system is what it was when we first exploited it: the most common case, and a workflow failure rather than an engineering one.

2. Fixed and regressed. The fix was real, then undone — a merge conflict resolved in favour of the older branch, a hotfix cut from a tag predating the fix, a rollback restoring a build that was known-good but not security-good, or an environment rebuilt from a template that still encoded the original misconfiguration. Configuration fixes are especially fragile: a setting applied by hand has a half-life measured in deployments, because the thing that creates the system does not know about it.

3. Fixed in one place only. The defect was a class and the report named an instance. That instance was fixed and verified. The same defect persisted on a sibling endpoint, a second build, a legacy API version still accepting traffic, or another environment never in scope. Nobody failed to fix the finding; they fixed exactly the finding, which is a reasonable reading of a ticket and an insufficient reading of a vulnerability. Scope boundaries do their quietest damage here: the tester stopped looking at the edge of scope, and so did the fix.

4. Compensating control removed. Nothing about the defect changed; the thing containing it did. The finding was accepted as a risk because a network restriction, an upstream filter, an authentication layer or a rate limit made exploitation impractical. That containment was later removed, relaxed or routed around — for a migration, a partner integration, or because the team owning it did not know what it was holding back. The acceptance was defensible when made, and never re-examined when its premise expired.

How to harden this

  • Separate the two states in your tracker. Fix deployed and remediation verified should be independent fields, set by different people. If one field carries both meanings, you cannot measure the gap.
  • Close findings on evidence, not on merge. Require the artefact type the tester produced: a request and response, a command and its output, a capture of the failed attack. "Fixed in release N" is a claim; a failed attack is evidence.
  • Fix the class, then enumerate the instances. Ask where else the pattern exists — other endpoints, builds, environments, API versions — and record the answer even when it is "nowhere else".
  • Encode configuration fixes where the system is built. A setting changed by hand will not survive a rebuild. Put it in the image, template or infrastructure definition, and the rebuild becomes an ally.
  • Add a regression test at the point of fix. A test that fails if the vulnerable behaviour returns turns a one-off fix into a standing assertion, re-checked for free.
  • Give every accepted risk an expiry and a named premise. Record what is containing the issue, and treat removal of that control as a change to the finding. A verified fix is valid as of a date and an environment, not forever.

Why 'not reproducible' is its own category

The most common dishonesty in retesting is not inventing fixes. It is quietly folding not reproducible into fixed. The two feel adjacent — in both cases the tester tried the attack and it did not work — and one is a clean green line in a report while the other is an awkward paragraph. But they are claims about different things. "Fixed" is a statement about the system; "not reproducible" is a statement about the attempt.

A finding can be unreachable for reasons that say nothing whatsoever about the defect:

  • Temporary environment state. A dependent service was down, a queue was not draining, a feature flag happened to be off, the component was mid-deployment. The vulnerable code path was not executing that week.
  • Changed test account. The original account no longer exists, or has a different role, or has lost an entitlement incidental to the vulnerability but essential to reaching it. The defect is intact; the key to the corridor is gone.
  • Missing data. Exploitation depended on records or relationships present a year ago and absent now. An access-control flaw over objects you hold none of is still an access-control flaw.
  • Changed surface, same defect. An endpoint renamed, a parameter restructured, a flow redesigned. The original steps fail verbatim; whether the weakness survived the reshuffle is a separate question that takes real work to answer.

None of these constitutes remediation. Reporting them as fixed is how a programme accumulates false confidence: the trend line improves while the attack surface does not, and the first to find the discrepancy is an attacker or next year's tester.

Handling the category honestly requires only saying something less tidy than "closed":

  • Report it under its own disposition, visible in the summary alongside still-present and fixed, never merged into either.
  • State the blocker precisely — which account, which data, which dependency, which state — so the reader can judge whether it is incidental or meaningful.
  • Say what would settle it: a specific access, a seeded record, a flag enabled in a test environment. Name it rather than leaving the reader to guess.
  • Keep the original severity. A not-reproducible finding is an open question at its original severity; downgrading it on the strength of a failed attempt is the same error in slower motion.
  • Carry it forward as unresolved, blocker recorded, so it is attempted first next time rather than rediscovered from scratch.
  • Where the blocker is removable within the engagement, remove it and finish the test. A retest ending in four not-reproducible findings because nobody asked for an account on day one is a scheduling failure dressed as a result.

Treated this way, "not reproducible" stops being an embarrassment and becomes useful: a dated record of what the retest could not establish. A report that admits uncertainty in four places is more trustworthy than one claiming certainty everywhere — and more useful to the engineer deciding what to do on Monday.

What "fixed" has to mean

Most of that decay is a failure of definition rather than of engineering. Teams that close findings quickly are not usually lying to themselves; they are accepting a signal that resembles verification and is not. So it is worth being precise about what verified remediation requires.

A finding is verified as fixed when four things are true at once:

  • The original reproduction steps are re-run — not an approximation, and not a scanner's equivalent check, but the same request sequence, role, object and preconditions, against the current build.
  • Someone who did not implement the fix runs them. The person who wrote the patch knows what they intended it to do, and that knowledge is precisely the bias being removed. Verification is an independence requirement before it is a skill requirement.
  • The result is captured as evidence. The denial, the empty result set, the absent header — something a third party can judge for themselves months later.
  • The verification is dated and tied to a specific build or environment. "Fixed" is not a property of a system but of a version of a system in a place. Without the version and the place, the claim cannot be falsified later, and so cannot be trusted later.

Against that, consider the signals teams accept instead. A closed ticket describes a workflow, not a system. A developer's assurance describes the change made rather than the behaviour that results, and that gap is where most recurrence lives. A configuration diff proves a setting changed in one definition; it says nothing about whether that definition is in force, survives the next deployment, or is ignored by a second path. A passing scan is the most seductive: it shows only that one tool, with one set of signatures, from one vantage point, no longer sees what it saw before — close to meaningless for an authorisation defect or a logic flaw.

These are reasonable triggers for verification. They are not substitutes for it.

Retesting the fix, not the symptom

When a tester reports an injection, a traversal or a missing authorisation check, the payload is not the vulnerability; it is the instrument that revealed it. The vulnerability is the absent check: input reaching an interpreter without being parameterised, a path used before canonicalisation, an object returned without asking whether the caller is entitled to it.

A fix that blocks the reported payload and nothing else leaves the defect in place: a denylist rejecting the exact string from the report, an edge rule written against the single captured request. The finding closes, the scanner goes quiet, and the defect waits for someone who types it slightly differently.

Retesting properly means re-deriving the finding rather than replaying it, varying along four axes:

  • The input. A different payload shape, length and position. If a numeric identifier was the vector, try another identifier, another format for it, and one belonging to a different owner.
  • The encoding and representation. Alternative and double encoding, mixed case, unicode equivalents, a JSON body where a form body was used, an array where a scalar was expected. Many "fixes" are string comparisons, and those are defeated by representation.
  • The method and transport. The same operation as a different verb; the parameter moved from body to query string or header. Per-route and per-verb controls are a recurring source of partial fixes.
  • The entry point. The same operation through another interface — the machine API rather than the browser flow, a bulk endpoint, an export function, a legacy version of the route still answering requests.

Then the question of where the control now lives. Enforcement at the edge — a gateway rule, a proxy filter, a client-side restriction — changes what reaches the application without changing what the application does. That is sometimes a legitimate interim measure, and should be recorded as one, but it is not a fix in the component that owns the decision. So reach that component in a way that bypasses the intermediary, and see whether the behaviour holds. An application only safe because something in front of it filters is not safe; it is shielded, and shields are configuration.

The general principle: a symptom patch rejects a known-bad input; a root-cause fix changes how the decision is made. So the retest asks not "does my old payload still work?" but "is the missing check now present, in the component that owns the data, for every representation of the request that reaches it?" If the fix cannot be described as a change in where and how a decision is taken — parameterisation rather than concatenation, authorisation against the object rather than the route, allowlist rather than denylist — it is probably a filter.

How to harden this

  • Require every remediation to be described by the control introduced, not the input blocked. If the description names a payload, send it back.
  • Enforce authorisation in the layer that owns the data, so every caller inherits it, rather than in each route handler.
  • Prefer structural defences — parameterisation, type-safe deserialisation, canonical identifiers, allowlists — over inspection of request content.
  • Record edge mitigations as temporary, leave the finding open until the owning component is fixed, and give the retester a build reachable directly rather than through the edge.

Where regressions hide

A finding is reported where the tester happened to find it. That place is an artefact of the engagement's time budget, not a boundary of the defect, and the surest way to produce a recurrence is to fix exactly what the report describes and nothing more. Where the same defect turns up again, roughly in order of frequency:

  • Sibling endpoints and parallel code paths. The create handler is fixed and the update handler is not; the detail view is fixed and the list, search and export views are not; a helper is copied, the copy corrected, the original still called elsewhere.
  • Other builds and platforms of the same product. A feature implemented separately for different clients, or a shared service where only the tested front end had its calls corrected.
  • Other environments and tenants. A fix applied to one deployment and not the others, or a tenant-specific configuration never brought into line: the control is correct for the tenant that was tested.
  • Infrastructure rebuilds. A hardening change made by hand, then erased when the host, cluster or image is rebuilt from a definition still carrying the old settings. The fix was real; it was not where the system gets its truth from.
  • Dependency updates. A behaviour removed by local patching or configuration, then restored by a version bump, an upstream default change, or a transitive dependency reintroducing a removed component.

The remediation question is therefore never "where was this reported?" but "where does this pattern exist?" That is a code-search and inventory question, usually answerable in an afternoon — far cheaper than the next engagement finding the sibling.

How to harden this

  • Search the codebase for the pattern and fix every instance, recording the ones you chose not to change.
  • Apply fixes in the shared layer where one exists, removing duplicated implementations rather than correcting each copy.
  • Reconcile every environment and tenant against the fixed configuration, and move hardening into the definition the environment is built from, so a rebuild reproduces the fix rather than reverting it.
  • Pin dependencies and keep a list of behaviours you deliberately switched off, to re-check after each upgrade.

Turning a retest into a regression test

All of that is manual, and manual verification decays at the rate the people and the memory decay. The durable answer is to convert each confirmed finding into an automated check that runs on every build, and the discipline there is to assert the outcome, not the status code. A test that accepts any non-200 response passes when the endpoint is removed, when the service is misconfigured, and when a new error path leaks the data inside a 500. A test that asserts the security property is far harder to satisfy by accident: a request for another tenant's object returns no field of that object's data; an unauthenticated privileged operation leaves the record unchanged when read back as an administrator.

verification record
  finding reference      <internal id>
  original steps         re-run unchanged
  verified by            <not the implementer>
  build / environment    <build identifier> / <environment>
  date of verification   <date>
  variants attempted     encoding, verb, alternate entry point
  control location       owning service, server-side
  evidence               request/response pair, stored
  outcome                fixed | partial | not fixed | mitigated at edge
  regression test        <test identifier>, runs on every build
regression assertion (shape, not syntax)
  given    an authenticated session for tenant A
  when     requesting an object owned by tenant B
  then     the response contains no field value from that object
    and    the response contains no reference permitting a follow-up read
    and    the audit trail records a denied access
  # asserting "not 200" would also pass if the route were simply deleted

Done consistently, this moves verification from an annual event to a continuous one, and gives the annual retest a better job: checking what cannot be automated, and confirming the automated checks still mean what they claimed.

Be honest about the limits. Chained attacks, business-logic abuse, flaws needing a particular sequence of human actions, and anything whose exploitability is a matter of judgement rather than observation all resist automation — as does the whole class of defects nobody has found yet. Flaky security tests also get disabled faster than functional ones, so a few assertions that truly encode the property beat many that approximate it.

Note. A regression test written from a finding is the cheapest security artefact available: the steps exist, the outcome is documented, the original failure is already evidenced. If a retest produces nothing else, it should produce these.

Making it part of the programme

The sequencing that works, roughly in adoption order:

  • Include the retest in the engagement. Bought separately, months later, verification competes with new work for budget and usually loses. Scoping it in also means the same testers verify, with the same context.
  • Re-verify after major releases and infrastructure change. These are the two events behind most of the recurrence in the proportions above. A targeted re-check costs a fraction of a full assessment and catches regressions while they are cheap.
  • Keep a register of previously confirmed findings. Not the ticket system — a durable record of what was confirmed, the steps, the evidence, the build and environment verified against, the date, and whether an automated check covers it. Without it, every engagement starts from zero and "did this stay fixed?" cannot even be asked.
  • Re-test a sample of old findings each year, weighted towards the categories that recur: authorisation, configuration, and anything fixed in one place when the pattern existed in several. Treat a recurrence in the sample as evidence about the population.

Underneath all of this is one change in posture. Remediation is not an event that concludes a finding; it is a claim about current behaviour, and such claims expire. A programme that measures itself by findings closed is measuring its own paperwork. One that measures findings still fixed, verified against a named build on a known date, is measuring something real.

PXL includes a remediation retest as standard on penetration tests. We re-run the original reproduction steps against the fixed build, confirm whether the control is enforced in the right place, and record what we find — including where a fix is partial, where it holds only at the edge, and where the pattern survives somewhere we never reported it.

How much of your last report is still unfixed?

Our penetration tests include a remediation retest as standard — we verify the fixes and confirm closure in writing.

Scope a test with retest