PXL Security LTD, Sofia, Bulgaria Offensive security since 2014[email protected]
MethodologyPenetration TestingInfrastructure

Disproving the CVE: when the version banner lies

By PXL Security23 September 202620 min read

A scanner that reports a version-based CVE has not found a vulnerability. It has produced a hypothesis: if this service is the build its banner claims, and if it is configured the way the advisory assumes, then a known weakness may be present. Testing that hypothesis — and being willing to conclude it is false — is one of the least glamorous and most valuable things an engagement can do. This first part covers why banners mislead, what believing them costs, and how to disprove a candidate safely.

A banner is a claim, not a diagnosis

The mechanism behind a version CVE is simpler than its presentation suggests. The scanner connects, reads whatever the service volunteers — a protocol greeting, a response header, an error page, a static asset's build string, a certificate subject — and normalises it into a product and a version. It looks that pair up in a table of advisories and emits every advisory whose affected-version range contains the parsed value. No request exercised the vulnerable code. Nothing was proved. The output is a join between a parsed string and a database, dressed in the severity colouring of whatever it matched.

That inference breaks in several ordinary ways.

  • Backported fixes that leave the version untouched. Long-term-support distributions exist so that operators can take security patches without taking behavioural change. Maintainers lift the upstream fix into the version they already ship and bump only the packaging revision. The upstream string — the one in the banner, and the one the scanner parses — stays exactly where it was. Patched and unpatched builds are indistinguishable from the network, and the scanner will insist on the advisory indefinitely.
  • Vendor and appliance builds. Appliance firmware often carries a component version that corresponds to no upstream release at all: the vendor has forked, pruned or selectively patched. The version is a lineage, not a code identity.
  • Features compiled out or never enabled. Many advisories live in optional code: a parser for a protocol extension, an authentication module, a scripting interpreter, a legacy compatibility handler. If that code was excluded at build time, or is present but unreachable because the feature is off, the version range is satisfied and the vulnerability still does not exist.
  • Compensating configuration. The vulnerable path may need an option that is disabled, a role the anonymous caller does not hold, a listener bound only to loopback, or an upstream component that strips the input the bug requires. The code is there; the route to it is not.
  • Deliberately false banners. Operators told to "fix version disclosure" sometimes do so by lying: a proxy rewrites the header, or a directive pins a fictional string. Banners understate versions as often as they overstate them, and the understating case is the dangerous one — it manufactures findings that lead teams to patch what is already patched.
Note. None of this makes scanners useless. Version matching is a cheap, high-recall way of generating candidates, and high recall is exactly what you want at the top of a funnel. The failure is treating the top of the funnel as the bottom.

What it costs to believe the scanner

A false Critical is often defended as prudent — better to over-report than to miss something. It is not free, and the costs land on the people least able to absorb them.

The first cost is unplanned change. A Critical on an internet-facing appliance triggers the emergency path: an out-of-hours window, an upgrade of a device that terminates remote access or mail flow, a rollback plan, people on a bridge call. Emergency change is where outages come from. If the finding was never real, the organisation has accepted real availability risk and real fatigue to remediate nothing.

The second cost is misallocated remediation. Teams have a fixed number of change windows and a fixed amount of engineering goodwill per quarter. Spending them on an unreachable advisory means not spending them on the authentication gap, the stale administrative account, or the mail-authentication record that is actually being abused. Over-reporting does not add capacity; it reassigns it, usually away from what matters.

The third cost compounds, and is the one worth naming to a board: erosion of trust. A team that has twice patched a device under emergency conditions and twice found nothing wrong learns a rule — that reports from this source are inflated — and applies it to the next report too, including the one that was true. Crying wolf is not neutral over-caution; it is the destruction of a signal you will need later. A report whose Criticals are always real can demand action. One whose Criticals are scanner output cannot, and eventually will not get it.

The inverse is rarely stated plainly: a disproof is a deliverable. "These seven candidates were probed with the real exploit primitive and none is exploitable" is a result a client can act on. It cancels seven change windows, closes seven tickets, and tells an operator something true about their own estate that no scanner could.

Testing the primitive, not the version

The method is one discipline applied per candidate. Read past the title to the mechanism and answer a single question: what primitive does this vulnerability actually require? A specific protocol verb. A memory read that returns adjacent heap. An unauthenticated handler that emits source. A reflection that reaches the browser as executable markup. Then test for that primitive, directly and minimally, and record what you observed. The version stops mattering the moment you have looked at the thing it was standing in for.

Four abstracted examples, from one external engagement where roughly seven Critical and High version-based candidates were each handled this way. Every one proved not exploitable.

An edge VPN appliance, memory-disclosure class. This class depends on a request whose declared length exceeds the data supplied, causing the service to return bytes it did not intend to send. That is a primitive you can ask for directly and observe without changing anything.

probe:       pre-auth request, declared length > supplied body
observable:  bytes returned vs bytes supplied
result:      returned length == supplied length
             no trailing heap content across repeated attempts
settles it:  the over-read primitive does not exist here

The observable is a repeated length comparison: if the leak exists, the response is longer than the input and its tail varies between attempts. Identical, exactly-sized responses mean the bounds check is in place, whatever the banner says.

A mail transfer agent, chunking remote-code-execution class. This class requires a specific transfer extension; the defective parser is reachable only through that extension's verb. The question is not "which version is this" but "will this service accept that verb at all".

> EHLO probe
< 250-SIZE …
< 250-STARTTLS
< 250 HELP
             # extension not advertised
> <chunked-transfer verb> 4
< 500 unrecognised command
settles it:  vulnerable parser is unreachable — verb neither
             advertised nor accepted, pre- and post-TLS

Note the two-step. Absence from the advertised extension list is suggestive, but some services accept unadvertised verbs. Issuing the verb and receiving a hard rejection — checked both before and after the TLS upgrade, since capability sets differ — is what closes it.

A web application, unauthenticated code-retrieval class. This class turns on a handler that returns application source to a caller with no session. We requested it with the retrieval parameter set and looked at the body, not just the status code.

probe:       unauthenticated request to the affected handler
             with the source-retrieval parameter present
observable:  response body
result:      404, generic error document, zero bytes of source;
             identical with and without a session
settles it:  the path is patched — handler absent, not merely
             access-controlled

A gateway interface, reflected cross-site-scripting class. Here the primitive is script execution in the victim's browser. We injected a canonical reflection probe and inspected both the reflected bytes and the response headers.

observable 1: reflected value is entity-encoded
              (< > " ' all escaped in the response body)
observable 2: Content-Security-Policy present, no inline
              execution permitted, no wildcard script source
settles it:   even if encoding were bypassed, policy blocks
              execution — two independent controls

Both matter because each answers a different question. Encoding tells you the injection never becomes markup; the policy tells you what would happen if it did. One control is a note about fragility; two independent controls is a disproof.

On that engagement, the risk that survived this process was entirely low-severity hygiene: weak email-authentication posture, version and information disclosure, and a legacy anonymous file-transfer service. A materially different — and far more actionable — picture than seven Criticals.

How to harden this

  • Record the packaging revision, not just the upstream version. Inventory from the package manager and firmware build identifiers rather than network banners, so backported fixes are visible to you even when a scanner cannot see them.
  • Keep a per-service feature manifest. Which optional modules, protocol extensions and interpreters are enabled? Most version advisories triage in minutes against an accurate manifest.
  • Require a primitive statement before opening an emergency change. If nobody can say which specific capability the exploit needs, nobody has triaged the candidate yet.
  • Do not remediate version disclosure by faking versions. Suppress the banner if you must; a false string corrupts your own triage as well as an attacker's reconnaissance.
  • Prefer defence in depth you can observe. Output encoding plus a restrictive content policy; unadvertised extensions plus rejected verbs. Independently verifiable controls are what turn candidates into disproofs.
  • Ask your testers for the observable. Findings and disproofs alike should state what was measured. "Version in range" is not an observable.

Safe disproof

Disproof has to be cheaper and safer than exploitation, or it will not happen on production systems — where it is most needed. A few rules carry most of the weight.

Prefer observation to exploitation. Most primitives announce themselves before they are used: a capability list, a response length, a header, a status code, the shape of an error. Observe first; escalate only if the observation is ambiguous.

Use the smallest primitive that separates the two worlds. The test should be the minimum interaction whose result differs between a vulnerable and a patched host. A four-byte chunk declaration establishes whether a verb is accepted; there is no reason to drive a full chain to find out.

Avoid destructive and state-changing proofs. No memory-corruption attempts against live services, no writes, no credential or session mutation, no restarts, no deliberately resource-exhausting requests. A disproof that takes down a mail gateway has cost more than the finding it cancelled.

Agree the approach where production is in scope. Say in advance which primitives you intend to probe, what each looks like on the wire, the worst realistic outcome, and when you will do it. Operators are rarely obstructive once they can see the plan is a length comparison rather than a shellcode attempt — and their logs will show the traffic, so they ought to recognise it.

Know when the honest answer is "not testable from here". This is the discipline that keeps the rest credible. If the primitive requires a session you were not given, a certificate you do not hold, a network position outside scope, or a state you may not create, you have not disproved anything — you failed to reach it. "Not vulnerable" means you tested the primitive and it was absent. "Not reachable from the tested position" means the candidate remains open and needs host-level confirmation of patch state or configuration. Collapsing the second into the first is the scanner's error, made by someone who should know better.

How you record each of these so a reader six months later can tell which you meant — and what residual risk survives a disproof — matters just as much as the probe itself.

Write the disproof down, or it didn't happen

A disproof that lives only in the tester's head is worth nothing. Not to the engineer who inherits the system next quarter, not to the auditor who re-runs the same scanner and sees the same red row, and not to the tester six months later, who can no longer remember which of the candidate issues on that host they actually chased down. The work was real; the absence of a record makes it unrecoverable.

The failure mode is predictable. At the next assessment, someone pulls the previous report, sees no entry for a flagged version-based issue, and has two hypotheses: it was tested and dismissed, or it was never looked at. Nothing distinguishes the two. So they either repeat the whole investigation at full cost, or — more often, under time pressure — copy the scanner's claim forward as an unverified Critical, and the organisation spends another cycle arguing about something that was settled a year ago. An undocumented disproof is operationally identical to an untested item. Treat it as one.

What makes an entry survive is that it records the reasoning, not just the verdict. "Not exploitable" is a conclusion with no load-bearing structure. Six things need to be on the page:

  • The candidate issue and its class. Not an identifier alone — the mechanism. A deserialisation flaw in a management interface, a path traversal in a static-file handler, an authentication bypass in a protocol handshake. The class is what a future reader reasons with; an identifier is just a lookup.
  • The primitive it depends on. What must be true for the published exploit to work: a module loaded, a feature flag on, a protocol version negotiable, a parser invoked on attacker-controlled input. This is the hinge of the entry.
  • The exact probe performed. Enough that someone else can repeat it verbatim: what was sent, to which interface, with which parameters, from which position in the network.
  • The observable response. What came back — the status, the error, the absent header, the rejected handshake, the silence. Observations, not interpretation.
  • The conclusion, with its scope. What the observation rules out, and only that. "The vulnerable parsing path is not reachable through the exposed interface" is a conclusion. "The host is secure" is not.
  • The conditions under which it stops holding. The expiry clause: if this flag is turned on, if this module is loaded, if this reverse proxy is removed, the disproof is void and the item returns to the queue.

On one external engagement we worked through roughly seven candidate Critical and High version-based findings. Every one failed to materialise: the component was present by version string but the primitive was missing in each case — a feature not compiled in, an interface bound to loopback, a parser never reached by the exposed request path, a configuration that disabled the vulnerable handler outright. Seven probes, seven recorded responses, seven conclusions with expiry conditions. Each entry took minutes to write and converted a week of future rework into a lookup.

The shape that holds up is boring and fixed. Boring is the point — a fixed shape is skimmable under pressure and hard to half-complete:

candidate   : remote code execution in <component class>, reported by version banner
class       : unsafe deserialisation of a request-body parameter
primitive   : optional serialisation module must be loaded AND the
              affected endpoint must be reachable from the test source
probe       : crafted request to the affected endpoint carrying a
              benign typed-object marker, from <test vantage point>
observed    : 404 on the endpoint; module-listing endpoint returns a set
              that does not include the serialisation module; no
              deserialisation error surfaced under malformed input
conclusion  : primitive absent — vulnerable code path not reachable
              through the exposed service. Not exploitable as deployed.
expires if  : the serialisation module is enabled; the affected endpoint
              is re-exposed; the service is reverse-proxied differently
residual    : component remains out of support window — tracked as
              patch-management hygiene
verified    : <engagement reference>, re-verify on next infra change

"Not exploitable today" is not the same as "not present"

The second discipline is resisting the relief. Having proved that a scanner's Critical does not fire, there is a strong pull towards deleting the row and moving on. That is the wrong call, for reasons that are entirely mundane.

The component is usually still outdated. What saved you was frequently a configuration — a module not loaded, an endpoint not exposed, a flag left at its safe default — and configurations get reverted by a rebuild, a template update, a rollback, or a colleague who needed that feature on Friday afternoon. Compensating controls get removed by people who do not know what they were compensating for. Features get enabled when an integration demands them. The vulnerable code is still on disk; the only thing that changed is that nobody is watching the condition that made it safe.

So the honest treatment is a downgrade, not a deletion. Report the residual as what it is: a hygiene or patch-management finding, typically Low, occasionally Medium where the protective condition is fragile or the component is well past end of support. State plainly that the exploitable path was tested and not found, and that the rating reflects the remaining exposure — unsupported software, a safety property resting on configuration rather than on fixed code. That framing is defensible to an auditor and useful to the team that owns the patch cycle.

On the engagement above, that is exactly where the residual risk landed: no exploitable remote code execution anywhere, and a short tail of low-severity items — weak email-authentication policy, version and internal information disclosure from banners and error pages, and a legacy file-transfer service still permitting anonymous access. None of those will get anybody on the front page. All are real, cheap to fix, and would have been invisible if the Criticals had been deleted along with the hypotheses that generated them.

How to harden this

  • Record the protective condition as a configuration assertion, not a memory — put the module list, bind address or feature flag under configuration management and alert on drift.
  • Keep the outdated component on the patch backlog with its real age, regardless of current exploitability.
  • Where safety depends on a proxy, firewall rule or access control in front of the component, document that dependency in the change-control record for the control itself, so removing it triggers a review.
  • Remove features and modules you do not use rather than relying on them being disabled; absent code cannot be re-enabled by accident.
  • Treat end-of-support dates as hard deadlines with owners, not advisory metadata.

The uncomfortable conversation about coverage

Sometimes you cannot reach the primitive at all — not because it is absent, but because you could not get near enough to find out. The service is filtered from your source address. Rate limiting makes a meaningful probe impossible. The interface sits behind an authentication layer you were given no credentials for. The window closed.

There is a comfortable thing to write in that situation and an honest thing. The comfortable thing is nothing at all, which quietly reads as "tested, clean". The honest thing is an explicit coverage limitation naming the asset, what could not be reached, why, and what would be needed to close the gap. An acknowledged gap is more valuable than a confident-sounding clean result, because a gap can be scheduled and a false clean cannot.

We learned the cost of this on a separate engagement. An entire address block in scope returned nothing at the application layer — no responses, no errors, no banners. The tempting read was that the range was unused or hardened to the point of silence. It was neither: the block was filtering at the application layer on the basis of our source address, and a second vantage point brought it to life. Had we not changed position, the report would have recorded a clean range that was in fact never examined.

Note. Silence is a result that needs explaining, not a result that needs recording. A range that answers nothing at all should prompt a change of source address, a change of protocol, or a line in the coverage section — ideally all three.

What a report full of disproofs is actually telling you

Buyers often mistake a report like this for a thin one. A scanner produced fourteen Criticals; the report presents two proven issues, seven documented disproofs and a handful of Lows. The instinct is that somebody went easy.

The opposite is true. Converting a long list of hypotheses into a short list of proven facts plus a defensible record of what was ruled out is the most expensive and most useful work in the engagement. Reproducing a known exploit is a morning's work; establishing that a vulnerable code path cannot be reached, and writing the evidence down so the conclusion survives the next rebuild, demands that the tester actually understand the deployment.

Read such a report by checking whether each dismissal is falsifiable. A credible disproof names the primitive and shows the probe; a thin test says "could not be confirmed" and moves on. The questions worth asking a supplier are blunt: for each item you dismissed, what had to be true for the exploit to work, and how did you establish it wasn't? Which conclusions depend on configuration that could change? What did you not reach, and from where did you try? Concrete answers mean you bought verification. Answers in the passive voice mean you bought a scan with a cover page.

Making this routine

None of this requires a consultancy. Internal teams can run the same discipline against their own scanner output, and the ones who do spend far less time arguing about severity labels.

Start by triaging by primitive rather than by severity. Group the queue by the condition each issue depends on — this module loaded, that endpoint exposed, this protocol version negotiable — and a long list collapses into a handful of questions about your own deployment, several of which clear whole clusters at once. Severity-first triage forces you to re-derive the same reachability question over and over in descending order of panic.

Then keep the register. One row per disproved item, in the shape above, with the expiry condition filled in, because the expiry condition is the part that earns its keep. Re-verify after infrastructure changes — a platform migration, a proxy replacement, a base-image update, a change to the ingress path — and treat a change touching a recorded protective condition as an automatic trigger rather than something to notice later.

Finally, align the process so a disproved item is tracked, not closed. Most vulnerability-management workflows offer only "open" and "resolved", and resolving something that is still installed loses the record. A third state — verified not exploitable, with conditions — keeps the item visible, keeps the residual hygiene finding on the patch backlog, and means the next person to see that scanner row finds an answer instead of a blank.

How to harden this

  • Add a "verified not exploitable (conditional)" state to the tracker, with mandatory fields for primitive, probe, observation and expiry condition.
  • Triage scanner output in primitive groups; answer the reachability question once per group, not once per finding.
  • Hook re-verification to change management: any change touching a recorded protective condition reopens the associated entries.
  • Make coverage limitations a first-class section in internal test records, so an unreachable asset is scheduled rather than assumed healthy.
  • Test from more than one network position wherever source-based filtering is plausible, and record the vantage point with every probe.
  • Review the register on a fixed cadence and promote anything whose expiry condition has quietly become true.

How many of your Criticals are real?

A validated vulnerability assessment separates the provable from the merely reported, so your team fixes what matters.

Request an assessment