PXL Security LTD, Sofia, Bulgaria Offensive security since 2014[email protected]
ComplianceBuyer's guidePenetration Testing

Does SOC 2 require a penetration test?

By PXL Security7 October 202619 min read

No. No Trust Services Criterion says "the entity shall perform a penetration test". If you came for that one fact, you have it. But "not mandated" and "not needed" are very different statements, and conflating them is how organisations end up in a difficult conversation with their auditor six weeks before a report deadline. Here is what the criteria actually say, and where the real obligation comes from.

The short answer

SOC 2 does not explicitly require a penetration test. The AICPA's Trust Services Criteria — the control criteria a SOC 2 examination is performed against — contain no criterion naming penetration testing as a mandatory activity. There is no line item you can fail for not having one: no minimum frequency, no required methodology, no stipulated scope.

Penetration testing appears in the criteria exactly once, and as an example — one of several illustrative evaluation types listed in a point of focus under criterion CC4.1. Points of focus are explanatory guidance, not requirements. We will return to that distinction, because it is the most misunderstood thing about SOC 2 and the source of most confusion here.

So far, so reassuring. Here is the part that matters more. SOC 2 is not a checklist standard but a set of outcome-oriented criteria evaluated by a licensed CPA firm exercising professional judgement. The question that determines your report is not "did you do the mandatory penetration test?" It is whether you can demonstrate, with evidence, that you have controls suitably designed and operating to detect vulnerabilities in your systems, and that management evaluates whether those controls are present and functioning. A penetration test is one of the most practical and auditor-legible ways to answer that, and for many organisations with an internet-facing product it is the only answer that holds up without a lot of compensating explanation.

In practice, then: the standard does not compel you. Your auditor's judgement, your risk assessment, your own control descriptions and your customers' questionnaires very often do. This part covers what the criteria genuinely say, so that when someone tells you a penetration test is "a SOC 2 requirement", you know precisely how wrong they are and in which direction.

Not legal or audit advice. PXL Security is an offensive security firm. We are not a CPA firm, not an auditor and not a certification body, and we do not issue SOC 2 reports or opinions of any kind. Only a licensed CPA firm can perform a SOC 2 examination, and only your service auditor can tell you what satisfies the criteria in your environment. Treat this as background for an informed conversation with them, not a substitute for it.

What the Trust Services Criteria actually say

To read SOC 2 accurately you need to know that the source document has two layers, and only one is binding.

Criteria are the benchmarks: the statements your controls are evaluated against, and your auditor must conclude on each applicable one. For the security category — the common criteria, present in every SOC 2 report — these are organised into nine groups, CC1 to CC9: control environment, communication and information, risk assessment, monitoring activities, control activities, logical and physical access, system operations, change management and risk mitigation. The first five are deliberately aligned with the COSO internal control framework and its seventeen principles; the remaining four address information and systems concerns COSO does not reach.

Points of focus are the second layer. Each criterion is accompanied by short statements illustrating characteristics a control meeting that criterion might have, to help organisations design controls and practitioners think about relevant evidence. They are explicitly not a checklist. The AICPA's guidance is clear that an assessment of whether each point of focus has been met is not required, that not every point of focus will be relevant to a given entity, and that they may be tailored or disregarded where they do not fit. The 2017 criteria remain in force with revised points of focus issued in 2022: the criteria themselves were not rewritten, only the illustrative guidance refreshed. That history is itself a clue about which layer carries weight.

With that structure in mind, here are the criteria that genuinely bear on security testing. We list only these four: claims that penetration testing is "required by" CC9.1, CC6.x or CC8.1 rest on inference about what a reasonable control might look like, not on criterion text.

  • CC4.1 — monitoring activities. The criterion requires that the entity selects, develops and performs ongoing and/or separate evaluations to ascertain whether the components of internal control are present and functioning. It maps to COSO Principle 16, and it is where penetration testing is named — not in the criterion text, but in a point of focus about considering different types of ongoing and separate evaluations, which offers penetration testing, independent certification against established specifications such as ISO, and internal audit assessments as examples. Note what the criterion is really about: monitoring. It asks whether management independently checks that its own control system works rather than taking its own word for it.
  • CC4.2 — deficiency communication. The entity evaluates and communicates internal control deficiencies in a timely manner to those responsible for corrective action, including senior management and the board where appropriate. This is why an unread test report can be worse than no test: findings that went nowhere are evidence against you.
  • CC7.1 — vulnerability detection. The criterion requires that, to meet its objectives, the entity uses detection and monitoring procedures to identify changes to configurations that introduce new vulnerabilities, and susceptibilities to newly discovered vulnerabilities. Its points of focus deal with defined configuration standards, change-detection mechanisms, detection of unknown or unauthorised components, and conducting vulnerability scans periodically and after significant change with timely remediation. Be precise here, because this is widely misquoted: CC7.1's illustrative guidance points at vulnerability scanning, not penetration testing. The two are not interchangeable, and treating a quarterly authenticated scan as though it discharged a penetration-testing expectation is a common and avoidable error.
  • CC3.2 — risk identification. The entity identifies risks to the achievement of its objectives and analyses them as a basis for determining how they should be managed. This is where a decision not to test has to be justified: if your risk assessment names application-layer compromise as significant and your control set contains nothing that would detect it, your auditor can reasonably probe the incoherence.

What this means for you

  • When a consultant, platform or vendor says a penetration test is "required by SOC 2", ask them to cite the criterion. If they cite CC4.1, ask whether they are quoting the criterion or a point of focus — the answer tells you how carefully they read the standard.
  • Do not conflate scanning and penetration testing in your control descriptions. If a control says "annual penetration testing", a scan report will not evidence it — your own wording is what the auditor tests against.
  • Decide which criterion your testing supports before you commission it. Our SOC 2 penetration testing page sets out how we scope against specific criteria rather than producing a generic report.
  • If you choose not to test, record why in your risk assessment, with compensating controls named. An undocumented absence is far harder to defend than a reasoned one.

Why the standard is written that way

The absence of a penetration-testing mandate is not an oversight. It is a design decision, and understanding it tells you how to behave.

SOC 2 is criteria-based rather than prescriptive. A prescriptive standard enumerates controls: do this, at this frequency, to this specification. PCI DSS works much more that way, which is why it carries an explicit penetration-testing requirement with stated scope and cadence while SOC 2 does not. A criteria-based standard states the outcome to be achieved and leaves the entity to select controls appropriate to its size, complexity, technology and risk profile. The rationale is practical: the criteria have to work for a SaaS platform, a payroll bureau and a data centre operator alike, and a control list specific enough for one would be absurd for another.

For a buyer, the consequence is uncomfortable but simple. Flexibility in the standard does not mean flexibility in practice. The standard declines to specify a control, so something else fills the vacuum: your auditor's judgement about what constitutes sufficient appropriate evidence, the controls you yourself asserted in your system description, prevailing practice among comparable firms, and your customers' expectations expressed through questionnaires and contracts. None of those is optional in the way the standard is. You have traded a clear rule for a negotiation — with people who have seen a great many SOC 2 engagements and hold a settled view of what normal looks like.

Type 1 and Type 2, and why it changes the answer

Which report you are pursuing changes the testing conversation substantially, and this is frequently missed.

A Type 1 report addresses the fairness of management's description of the system and the suitability of the design of controls as at a specified date. It is a point-in-time assessment of design: the auditor asks whether the controls, if they operated as described, would be capable of achieving the criteria. It does not test whether they actually ran. A Type 2 covers the same ground and adds operating effectiveness over a defined review period, commonly three to twelve months, which means sampling evidence dated inside that window.

The implication is direct. For a Type 1, a documented and approved penetration-testing policy with defined scope, cadence, remediation workflow and ownership can be enough to evidence that the control is designed. A test may not yet have happened; the control exists on paper, and paper is what a design assessment examines. Plenty of Type 1 reports are issued on that basis, quite properly.

For a Type 2, the policy is not evidence of operation — it is evidence of intent. If your system description asserts that the entity performs penetration testing, your auditor will look for a test actually performed, with dated artefacts, within the review period, plus evidence that findings were triaged and remediated or formally accepted. A test conducted two months before the period began does not demonstrate operation during the period. Neither does a report with open critical findings and no remediation trail: that evidences a control which ran and then failed at CC4.2.

Type 2 is therefore where the question stops being academic. A Type 1 can often be satisfied by good documentation; a Type 2 asks for proof that things happened inside a specified window, and testing is among the control activities most often found wanting when the evidence is pulled.

What this means for you

  • Establish which report you are pursuing before you plan testing. A Type 1 may be satisfiable with a well-drafted policy; a Type 2 wants artefacts dated inside the review period.
  • If you are moving from Type 1 to Type 2 — the usual path — assume every control you merely documented now needs operational evidence. Testing is a predictable item on that list.
  • Write control descriptions to match what you will actually do. Over-promising — "quarterly penetration testing" when you intend one test a year — creates an exception out of nothing.
  • Ask your auditor early what they expect to see for CC4.1 and CC7.1 in your environment. They are the only party who can answer that, and they would rather tell you in month one than during fieldwork.

Why you will probably be asked for one anyway

The absence of a mandate is not the absence of an expectation. In practice most organisations going through a SOC 2 examination commission a penetration test, for three reasons unrelated to any written rule.

The first is the service auditor's judgement about evidence. The criteria require management to evaluate whether controls are present and functioning, and require detection procedures capable of identifying new vulnerabilities introduced by configuration changes. How you satisfy that is your choice; the controls are not prescribed. But the points of focus published alongside the criteria do name penetration testing as one kind of evaluation management might use, alongside independent certifications and internal audit work. That is explanatory rather than mandatory, yet it shapes what an experienced auditor expects when they ask how you evaluate your own controls. If the answer is a quarterly automated scan and nothing else, you are asking them to accept a narrower evidence base than most of their clients present. Some will; many will push back.

The second is commercial, and it usually arrives before the audit does. Enterprise procurement and third-party risk functions work from standardised security questionnaires, and those ask about penetration testing directly, typically with a frequency and scope attached. The same requirement often appears in the security schedule of a master services agreement: an annual test of production by a qualified independent party, with a summary available on request. Once that clause is signed you are testing because the contract says so, and the SOC 2 report becomes a convenient place to evidence it.

The third is your own risk management, and it is the only one actually about security. A SOC 2 report describes a system and asserts that controls over it operated effectively. If you have never had a competent adversarial assessment of the authentication paths, the tenant isolation logic or the cloud control plane behind that system, you are asserting something you have not tested. Scanners find known defects in known software; they do not reason about broken access control between two of your customers, or a privilege escalation path that exists only because of how your IAM roles are wired together. In a multi-tenant platform those findings matter most, and no tool will surface them.

None of this makes a test mandatory. It makes it the path of least resistance — practice and commercial reality rather than a rule.

When to test within the audit period

For a Type 2 report, timing is the decision that most often determines whether the test helps or merely costs money. A Type 2 opinion addresses whether controls operated effectively throughout a specified period, so everything you want that report to say about your testing has to have happened inside that window.

The sequence is straightforward. Test early — the first third of the observation window is a reasonable target. That leaves time to triage findings, remediate the ones that warrant it, document a risk decision on the rest, and have the tester confirm closure. All of it then sits inside the period, and the auditor can examine the whole control cycle rather than just the test: identification, assessment, action, verification. That cycle is the evidence; the report alone is a quarter of it.

Leave it late and the arithmetic turns against you. A test commissioned six weeks before period end produces findings you have not yet fixed when fieldwork begins. You can remediate quickly, but a retest is unlikely to be completed and documented before the window closes — leaving a dated record of known weaknesses in the system the report describes, with nothing after it. That is a harder conversation than having had no test at all, and a plausible route to an exception if the control you described included timely remediation.

A test completed after the period ends adds very little to that report. It did not happen during the period under examination, so it cannot evidence a control operating during it. Teams sometimes expect a late test to be credited retrospectively because little has changed; auditors are generally unwilling, and the reasoning is sound — a test assesses a specific build at a point in time, and its relevance decays as the environment moves on. First-time periods are often short, three months being the shortest commonly seen in practice, leaving almost no room at all.

What this means for you

  • Book the test for the first third of the observation window, not the last.
  • Budget calendar time for remediation and retest inside the period — how much depends on your release cadence and the severity mix, so agree it with your testers rather than assuming.
  • If the period has started and you have not tested, do it now rather than waiting for a convenient release boundary.
  • Assume nothing completed after period end will count towards that report.

What the auditor actually wants to see

Auditors do not assess the technical quality of your test; that is not their role. They examine whether a control you described operated as described, which for a testing control resolves into a small set of artefacts.

  • A defined scope that ties to the system description. The report's scope section should name the environments, applications, domains and infrastructure assessed in terms a reader can map onto the described boundary, with dates, the authenticated roles tested, and any exclusions and their reasons.
  • Evidence the work was done by someone suitably independent and competent. The tester need not be external; a skilled internal team independent of the builders can be acceptable. Auditors look for some basis for believing the testers were qualified and were not assessing their own work — a named firm, a stated methodology, credentials, or in-house documentation of segregation and skills.
  • The findings, with severities. Each rated under a stated scheme, with enough detail for a remediation owner to act. Severity matters because your own policy ties remediation timescales to it, and that policy is the yardstick your response is measured against.
  • Evidence of remediation. Tickets, change records, commit references, timestamps — showing findings were assigned, actioned and closed within the timeframes your policy commits to. Where you chose not to remediate, a documented risk acceptance with a named approver and a rationale serves the same purpose.
  • Confirmation of closure. A retest is the cleanest form, ideally a separate dated document stating which findings were verified fixed and which were not. Where a full retest is disproportionate, targeted verification evidence can serve, provided it demonstrably addresses the specific finding.

A PDF with no remediation trail is weak evidence. It shows you bought a test; it does not show that your detection, assessment and remediation process worked, which is the control you described. The strongest package is a test report, a remediation log referencing its finding identifiers, and a retest letter — three documents that read as one continuous story. Our sample report shows the scope statement, finding format and retest section.

Your auditor decides what they accept. Nothing above is a rule. Service auditors exercise professional judgement about the sufficiency of evidence, and firms differ on what qualifies as a testing control, how they treat internal testers and what counts as confirmation of closure. Agree the expected artefacts with your audit firm before you commission the test, not after you receive it.

Scoping a test that supports the report

Scoping a test to support a SOC 2 report is not a technical question first. It is a question about the boundary described in that report — the infrastructure, software, procedures and data used to deliver the services in scope. Start from the system description, list the components inside the boundary, and work out which can be meaningfully assessed from an attacker's position.

For a typical SaaS platform it resolves into four areas. The production environment, or a staging environment demonstrably built from the same code and configuration, since findings against a dissimilar environment say little about the system under examination. The authentication and session paths — login, password reset, multi-factor enrolment and bypass, SSO, token and API credential handling — the gate in front of everything the report asserts. The multi-tenant boundaries, tested with at least two accounts under separate tenants so access control between customers is exercised rather than assumed. And the supporting cloud infrastructure: IAM roles and trust policies, storage exposure, network segmentation, secrets handling, container configuration, and the deployment pipeline where it reaches production.

The common mismatch is mundane and expensive. An organisation tests its public website, because that was the obvious external surface, while the SOC 2 report covers the customer platform behind the login. The test is valid and entirely irrelevant to the report: nothing inside the described boundary was assessed. Subtler versions abound — the web application but not the API serving the same data, one region while the description covers three, a pre-production tenant without the production authorisation logic. The auditor reads the scope section and sees the gap; the remedy is usually a second test on a compressed timeline.

Avoiding it takes one meeting. Put the system description and the proposed scope side by side, and make somebody account for every in-boundary component not being tested. Scoping for compliance-driven penetration testing should begin with that document, not a list of URLs.

What this means for you

  • Derive the scope from the SOC 2 system description, not last year's scope or whatever is easiest to reach.
  • Include the authenticated application and its API, not just the unauthenticated surface.
  • Provision at least two tenants and the full range of roles so isolation and privilege boundaries are genuinely tested.
  • Cover the cloud control plane and the deployment pipeline if they sit inside the described boundary.
  • Write down, and have someone approve, the reason for every in-boundary exclusion.

Mistakes that cost time at audit

The failures below reliably create extra work during fieldwork. None is technical; all are planning failures.

  • Presenting a vulnerability scan as a penetration test. Scanning is valuable and should run continuously, but a tool-generated findings list has different limits: it will not exercise tenant isolation, business logic or chained privilege escalation. Auditors who read many of these spot the difference from the finding titles alone.
  • Testing outside the period. Still the most common timing error. A test dated two weeks after period end evidences nothing about that period.
  • No retest or confirmation of closure. Remediation that was never verified is an assertion, not evidence — the easiest gap to close and the one most often left open, usually because the retest was not in the original budget.
  • Scope that does not match the system description. Either the test covers less than the boundary or it covers something adjacent to it. Both leave the auditor unable to connect the evidence to the system under examination.
  • Findings left open with no documented risk decision. It is legitimate to decide a medium-severity finding will not be fixed. It is not legitimate to leave it there with no record of who decided, on what basis, and when the position will be revisited. Undocumented inaction looks identical to neglect.
  • Assuming last year's test carries over. A test from the previous period evidences the previous period. If your control commits to annual testing, each period needs its own test, remediation trail and verification. Bridge letters speak to control continuity between periods; they do not substitute for testing inside the current one.

Need a test your SOC 2 auditor will accept?

We scope the test to your system description and deliver the evidence trail — findings, remediation guidance and a retest letter.

Scope a compliance test