PXL Security LTD, Sofia, Bulgaria Offensive security since 2014[email protected]
Buyer's guidePenetration TestingMethodology

What drives the cost of a penetration test

By PXL Security4 October 202620 min read

A penetration test is not bought by the unit. It is skilled attention applied to a system for a bounded period, so the price follows from one question: how long does this take to do properly? Every factor below lengthens or shortens that period, and knowing which ones you control is the difference between buying an assessment and buying a document.

Scope size is the obvious driver, and the least interesting

Scoping conversations begin with counts, because counts are the part a buyer can answer unaided. How many applications, APIs and endpoints. How many external hosts, internal subnets and live addresses within them. How many cloud accounts hold production workloads. How many mobile builds — one codebase shipped twice, or two separate applications?

Counts are a floor — someone must enumerate and sanity-check every item named — but they predict effort poorly, because effort tracks distinct behaviour, not inventory. Consider ten REST endpoints on one service: same framework, same authentication middleware, same authorisation model, same object-reference pattern. A tester establishes how that service behaves once; each further endpoint is then a confirmation pass — does it follow the pattern, and where does it deviate? Deviations are where findings live, but that is a fraction of the first endpoint's cost. Now two genuinely different applications: a server-rendered legacy system with session cookies and a bespoke permissions table, and a single-page application on a GraphQL API with token auth and federated identity. Nothing transfers between them: two mental models, two authorisation matrices, two sets of framework-specific weaknesses. Those two cost considerably more than the ten.

Infrastructure behaves the same way: a subnet of identical workstation builds is cheaper per host than a small one holding a domain controller, a build server and an unaccounted-for appliance. So the useful scoping question is not "how many?" but "how many different?" Telling a tester that forty endpoints share one gateway and one authorisation model while three are bespoke usually returns a smaller estimate than a flat count of forty-three.

Complexity is what actually costs

The real multipliers are structural: the properties that make a nominally small application expensive.

The number of distinct roles and permission levels is the most underestimated of them. Access-control testing is not a feature you check; it is a matrix you traverse. For each sensitive operation, a tester must establish which roles may perform it, then attempt it as each role that may not — vertical access control. Separately, for each role, they must attempt to reach data belonging to another instance of that same role: another user, another team, another tenant — horizontal access control. Two different tests, each applying to every operation that matters.

The work therefore grows with roles multiplied by operations, not roles added to operations. One extra role does not add one more check; it adds a full pass over every privileged operation, a new set of role pairings to probe in both directions, and an account with seeded data to support it.

Illustrative shapes, not estimates.

  R = distinct roles      O = sensitive operations

  vertical checks   ≈  R × O
  role pairings     ≈  R × (R − 1) / 2   (probed both ways)
  horizontal checks ≈  O × objects, per tenant boundary

  R = 2   user, admin
          vertical → 2 × O      pairings → 1
  R = 6   anon, user, team-admin, org-admin, support, auditor
          vertical → 6 × O      pairings → 15

  adding one role at R = 6:
          + O vertical attempts
          + 6 new pairings
          + another account and data set to maintain

Three roles is a manageable matrix. Seven, with overlapping grants, per-object permissions and a "support may act as user" capability, cannot be hurried without simply not doing most of it. A quote treating a seven-role application like a two-role one has decided to sample rather than traverse.

Bespoke logic versus standard CRUD. Records handled through a conventional framework have a well-understood weakness surface. Logic that prices things, moves money or enforces regulatory rules exists only in your application and must be learned first. Such flaws are found by understanding intent and then violating it — skipping a step, replaying one, supplying a value the designer assumed impossible — and that comprehension is not free.

Multi-tenancy. Isolation makes every data-access question two-sided and needs at least two provisioned tenants with distinguishable data. Enforced consistently in one place it is far cheaper to assess than when enforced per query, per service or per cache key.

Unusual authentication. Custom token formats, home-grown single sign-on, federation across several providers, step-up authentication, device binding and signed client-side state each require analysis before testing begins. Well-trodden authentication does not.

Heavy state and workflow. Where outcomes depend on sequence — onboarding stages, document lifecycles, case management — testers must reach each state before testing it, and rebuild it after every destructive attempt. That setup time is invisible in a scope document and very visible in a schedule.

Third-party integrations. Each adds a trust boundary and usually a callback worth attacking. Those exercisable only against a live third party often cannot be fully tested — a scoping decision, not a saving.

What this means for you

  • Describe distinctness, not just counts. Say which assets share a framework, gateway and auth model; grouping honestly usually reduces the estimate.
  • Bring a role list, flagging roles that are cosmetic variations on the same grants. Those can often be covered together.
  • Prioritise, and phase large estates. Crown jewels thoroughly now, long tail next, beats uniform thinness.
  • Do not trim roles or tenants to save money. Removing a role removes the test, not the attack path. Reduce breadth instead, and record what you excluded.

Depth: what level of assurance are you buying

Two quotes for the same scope can differ enormously and both be fair, because they are offers to do different things. Depth is a ladder, and each rung answers a different question.

Automated scanning with human validation. Tooling tests for known issues; a human triages and confirms what is real. It answers: do we have known, tool-detectable weaknesses? It will not find logic flaws, broken access control between roles, or anything needing an understanding of what the application is for.

Time-boxed manual testing. A tester works a methodology against the scope within an agreed period, pursuing what looks interesting and stopping when the time is spent. It answers: what would a competent attacker find within a bounded effort? Coverage is methodical but finite, and this is what most people mean by a penetration test.

Objective-driven testing. Defined by goals rather than coverage: reach customer data, obtain administrative control, cause a transaction to settle incorrectly. It answers: can a specific consequence be achieved? Coverage is uneven by design, and the finding count may be low while the result is severe.

Adversary simulation. An exercise against the organisation rather than the asset, usually assessing detection and response, often under stealth requirements. It answers: would we notice, and could we stop it? Planning and deconfliction overhead come with it.

These are not discounts of one another. A scan is not a cheap penetration test and a red team is not an expensive one; they are different products with different outputs, evidence and failure modes, so comparing their prices is a category error. The commonest way buyers end up disappointed is buying one rung while expecting the answers of another. Decide which question you need answered, then compare only like with like.

Access and environment

How much the tester is given changes both cost and value, and the relationship is the opposite of what many buyers assume. Black box means minimal prior knowledge: a name, a range, a URL. Grey box adds credentials for each role, documentation, architecture notes, API specifications and a named contact. White box adds source code, configuration and infrastructure definitions.

The instinct is that black box is cheaper, because less changes hands, and more realistic, because that is the attacker's position. Both are mistaken for most assessments: withholding information does not reduce work, it relocates it. A tester without credentials spends time inferring what you could have handed over on the first morning — enumerating roles, reverse-engineering a token format, establishing which service actually enforces authorisation. That time is deducted from testing, not added to the budget.

Realism is a weaker argument than it sounds, too: a real attacker is not time-boxed and may spend months on reconnaissance you have funded a fraction of, so simulating their ignorance inside a bounded engagement mostly simulates their first week. Where that is the actual objective — testing detection, or confirming an asset is not discoverable — black box is the right choice, deliberately. As a default it is a poor trade. Source access pushes the same way, turning "this behaves oddly, is it exploitable?" into a question answered by reading.

Environment matters as much as access. Production is the system that actually exists, and carries real limits on destructive activity. Staging permits more aggression but is only as useful as its fidelity: differences in authentication, infrastructure, feature flags or integrations make findings arguable both ways — issues that do not apply in production, and production issues invisible in staging.

Data quality is quietly decisive. An empty environment is close to untestable for access control: with no records to reach and no second tenant, there is nothing to attempt unauthorised access to. Then readiness on day one: accounts working at the right privilege level, enrolment completed, no lockout from rate limiting, someone reachable when something breaks. An engagement that spends its opening stretch chasing access never gets that stretch back.

What this means for you

  • Default to grey box, and go further where you can. Hand over credentials, documentation and specifications unless a stated objective requires otherwise.
  • Provide one working account per role, plus two tenants — verified by someone who has actually logged in with them.
  • Seed realistic data: populated, distinguishable records per account and tenant.
  • Clear platform obstacles in advance — rate limiting, bot protection, geo-blocking, short session timeouts — and document staging-to-production differences up front.
  • Name a technical contact who can act, not relay: someone who can reset an account the same day.

Constraints that add cost without adding value

This last category is the one buyers control most directly and examine least. None of it answers an additional security question.

Restricted testing windows. Testing is cumulative: a tester holds a model of the system in their head and loses time rebuilding context after every interruption. A narrow daily window does not deliver a proportionate fraction of the testing. It delivers less, because the overhead recurs.

Out-of-hours and weekend work. Sometimes genuinely necessary, for systems that cannot tolerate disruption during trading or clinical hours. It is also dearer per unit of work, and worth confirming the requirement is real rather than inherited.

Onsite presence. Occasionally mandatory — physical assessments, segregated networks — but often habit. It adds travel, reduces working hours, and for most remote-testable scopes improves nothing.

Heavy change control. Where every tool, source address and category of activity needs individual approval, the engagement proceeds at the speed of that process. Pre-approving the activity as a whole is far cheaper than approving it piecemeal while a tester waits.

Unstable environments. One that falls over under ordinary testing, or is simultaneously hosting a release, costs time twice: once in the outage, once in distinguishing a genuine finding from a broken deployment.

Slow access provisioning. The commonest of all — credentials arriving late, with the wrong privileges, or for the wrong environment. Unlike the others, it has no upside whatsoever.

Some of these will be non-negotiable in your organisation. The point is that each is a choice with a price paid in testing time, and you are paying it rather than being charged it. Remove two or three avoidable constraints before going to market and the same budget buys a materially deeper test.

Comparing quotes that aren't comparable

The buyer's problem is not that providers price differently. It is that "a penetration test" names a deliverable rather than a piece of work, so two proposals can promise the same thing, cost very differently, and both be honest. The difference sits in what will actually be done — usually the part the proposal says least about.

The variables that move effort most, in rough order of impact:

  • Depth. One engagement enumerates the attack surface and tests each class of weakness once. Another pursues chains: the low-severity disclosure that supplies an identifier, the identifier that reaches an endpoint missing an ownership check, the access that reveals an administrative function. Both are penetration tests. The second takes longer and finds the things that matter.
  • Seniority. Configuration weaknesses and known vulnerability classes are within reach of a competent junior tester. Business-logic abuse, authorisation modelling across roles, and anything requiring a hypothesis about how the system was built are not.
  • Coverage. One quote covers a single role, the primary journey and the documented API. Another covers every role, the undocumented endpoints the client reveals, the administrative interface and the integrations. Both scope statements may read identically.
  • Retest. Included as standard, sold as an extra, or absent — a material difference in what you receive, and frequently a footnote.
  • Reporting. A finding with reproduction steps, evidence of impact and specific remediation is a work item. The same finding as a severity label and a generic recommendation is a research task you have been handed.

So normalise before you compare. Take the cheapest proposal and write down what it commits to doing in concrete terms: roles tested, interfaces in scope, manual or tool-assisted, retest or not, evidence standard. Read each other proposal against that list and mark where it offers more. The residue — the part not explained by depth, seniority, coverage or follow-through — is the only part worth negotiating over. And treat vagueness as information: ambiguity in a scope statement is rarely resolved in the buyer's favour once delivery is under way.

The questions that make quotes comparable

Send the same questions to every provider and compare the answers rather than the proposals. These separate otherwise identical offers.

  • Who performs the test, and what is their experience? Not the company's credentials — the individuals'. Ask also whether a senior reviews the findings of anyone less experienced. A provider that will not name a level of seniority is telling you it may not be available.
  • Is this manual testing, validated scanning, or both? Tools belong in a good engagement; they cover known classes quickly and free time for the parts needing judgement. The question is what proportion of the effort is human.
  • How many roles will be tested, and how much will genuinely be covered? Ask for the answer in roles, interfaces and functional areas. A provider who knows what they are quoting can say what will be reached and what will not.
  • What does the report contain? Ask for a redacted example — the only reliable way to judge reporting quality before delivery. Ours is published as our sample report, and any serious provider will have something equivalent.
  • Is a remediation retest included? If so, what does it cover and how long after delivery can it be used? If not, what does one cost — and compare that against the quotes that include it.
  • Will findings be demonstrated with reproduction steps? A finding your engineers can reproduce is one they can fix and later verify. One they cannot reproduce becomes an argument about whether it is real.
  • What happens if something critical is found mid-test? You want immediate notification, not discovery at the report stage. Ask who contacts whom, how quickly, and whether testing pauses.
  • Will the same people debrief our engineers? A conversation with the person who found the issue resolves more than a written report does. If the debrief comes from an account manager, much of the value stays with the tester.

What this means for you

  • Send one list of questions to every provider and compare the answers side by side, not the proposal documents.
  • Judge each sample report as an engineer would: could your team act on this finding without asking a question?
  • Get the retest position in writing — what it covers and the window in which you can use it.
  • Treat an unanswered question as an answer, and keep the written responses as an annex to the contract.

What to give the tester to get an accurate quote

Quotes are only as good as the information behind them. A provider working from a one-line request is estimating against an imagined system, and that estimate will be wrong in one direction or the other: too low, and the scope quietly narrows during delivery; too high, and you pay for uncertainty. A short scoping pack removes most of that, and it need not be polished:

  • An architecture description of a paragraph or two: what the system does, the major components, what talks to what, where the data lives.
  • Counts. Applications, APIs and roughly how many endpoints, distinct user roles, environments.
  • Technologies. Languages, frameworks, cloud platform, authentication mechanism, anything unusual. This determines who should be assigned to the work.
  • Whether production or staging is in scope, and how faithfully staging reflects production.
  • Any compliance driver and deadline. A certification, customer requirement or contractual obligation shapes the reporting and the evidence; the deadline shapes the sequencing.
  • Known constraints. Change freezes, testing windows, rate limiting, third-party components you cannot authorise, anything known to be broken.

A provider who quotes without asking for any of this is guessing, and you will meet the consequences during delivery. Conversely, a scoping call that works through your architecture in detail is not a sales technique — it is the work of deciding what to test and who should test it. Scoping is the engagement's first technical judgement, and a provider who takes it seriously is telling you something.

Note. If a scoping conversation tells you something about your own system you had not written down, that is an early signal about the provider. If it is purely commercial, that is a signal too.

What a cheap quote is usually buying

Cheap is not automatically bad: a small, well-defined scope tested properly can be inexpensive. But a quote substantially below the others is buying something, usually one of these — and each is detectable from the proposal rather than after delivery.

  • Automated scanning presented as testing. In the sample report, look for findings that could only have come from a tool: generic descriptions, no evidence of impact, severity lifted straight from a scoring table.
  • Junior testers with no senior review. Ask for named individuals, their experience, and whether a senior reviews findings. A refusal to discuss the team is the answer.
  • No retest. Read the deliverables list rather than the narrative. If a retest is not listed as included, assume it is not, and price one in before comparing.
  • A templated report with no reproduction steps. Visible in the sample: findings written against a generic application rather than yours, and remediations that merely restate the vulnerability class.
  • Offshore or subcontracted delivery you were not told about. Ask where the testers are, who employs them, and whether any part of the work is subcontracted. Nothing is wrong with a distributed team; something is wrong with not being told.
  • A scope quietly narrowed to fit the price. Compare scope wording across proposals. Phrases such as "representative sample", "key functionality" or "as time allows" transfer the coverage decision to the provider after you have signed.

What this means for you

  • Read the deliverables list and scope definition before the methodology narrative. That is where the commitments and the omissions are.
  • Challenge every elastic phrase in a scope statement and ask for a count instead — roles, interfaces, functional areas.
  • Ask where the work will be performed, by whom, and whether any of it is subcontracted.
  • Normalise the cheap quote upwards by adding the retest and the coverage it omits, then compare again.
  • If a provider cannot produce a sample report, assume the report is a template.

Where spending more is and isn't worth it

Not every increase in price buys a proportionate increase in value. Worth paying for:

  • Senior testers on complex logic. Wherever correctness depends on business rules — entitlements, workflow states, pricing, approvals, tenant separation — experience determines what is found. This is the least substitutable thing you can buy.
  • Depth on the systems that matter. Concentrated effort against the application holding the sensitive data, or the component whose compromise affects every customer, returns more than the same effort spread thinly.
  • A retest. Without verification you have a list of claimed fixes. The retest converts remediation effort into evidence, cheaply relative to the work it validates.
  • A report your engineers can act on. Reproduction steps, evidence and specific remediation remove the time otherwise spent reconstructing the finding — a cost you pay internally when the report is thin.

Often not worth paying for:

  • Testing everything at equal depth. Uniform coverage sounds rigorous and reliably underweights the systems that matter. Most things deserve a look; a few deserve sustained attention.
  • Repeating a broad annual test when a targeted one would answer the real question. If what you need to know is whether the new payment flow is sound, scope that and test it properly rather than re-covering unchanged ground.
  • Scanning you could run yourself. Vulnerability and dependency scanning belong in your pipeline, running continuously. Buying them as a once-a-year service at consultancy rates is poor value in both directions.
  • Adversary simulation before the known issues are fixed. Red teaming answers whether you would detect and respond to an intrusion. If an unauthenticated interface is still exposed, you already know the answer, and a penetration testing engagement will tell you more per unit of effort. Simulation earns its cost once the obvious paths are shut.

Getting value from the test you buy

Once the engagement is agreed, how much you get from it is largely in your hands. Most wasted effort in a test is administrative, not technical.

  • Have accounts and environments ready before the start date. Credentials for every role, confirmed working. Testing time spent waiting for access is time you have bought and not used — the most common avoidable loss there is.
  • Give the testers a named contact. Someone who knows the system and can confirm whether a behaviour is intended. A tester blocked on an unanswered question either stops or guesses, and both are expensive.
  • Plan remediation capacity before the report lands. A report delivered into a fully committed sprint sits until the next quarter, and the risk persists for exactly as long. Reserve the engineering time when you book the test.
  • Use the retest. Fix, then have the fixes verified against the original reproduction steps. Unverified remediation has a habit of being partial — a control added at the edge rather than in the component that owns the decision, or applied only to the reported endpoint.
  • Keep the findings. The reproduction steps are the cheapest regression tests you will be given.

PXL includes a remediation retest as standard on penetration tests. We re-run the original reproduction steps against the fixed build and record what we find, including where a fix is partial — because a finding closed without verification is a claim, while one verified against a named build is a fact.

Want a scoped, itemised proposal?

Tell us what you're running and why you're testing. A senior tester replies with questions, a proposed approach and a quote.

Request a proposal