PCI DSS Compliance Testing How to Plan and Pass in 2026

Only 27.9% of assessed organizations achieved 100% PCI DSS compliance during interim validation in 2019, according to Verizon’s 2020 Payment Security Report. That figure changes how I approach PCI DSS compliance testing. The hard part isn’t scheduling a scan or producing a penetration-testing report. It’s proving, repeatedly and with current evidence, that payment systems, segmentation boundaries, applications, identities, logs, and operational controls still work as intended after the environment changes.
A credible program treats testing as a living control-validation process. It connects scope discovery to test design, test results to remediation tickets, and remediation to retesting. It also uses safe methods for checking real application behavior, including masked production-traffic replay in isolated environments, instead of assuming that an annual scan captures every path into a modern cardholder data environment.
Why PCI DSS Compliance Testing Fails and What Changed in 2026
Organizations fail testing for predictable reasons. They define scope from an old diagram, rely on firewall intent rather than observed behavior, collect reports only when an assessment is approaching, and treat remediation as a documentation exercise. Mature teams can still drift because cloud resources, payment-page dependencies, identity systems, build pipelines, and third-party connections change faster than formal assessment cycles.
The compliance history explains why annual-checklist thinking no longer works. PCI DSS was first released in December 2004, and the PCI Security Standards Council later marked September 2006 as the first release in its 20-year timeline. Major revisions followed, including PCI DSS v2.0 in October 2010 and v4.0 in March 2022, which the Council described as the biggest update since the original release. The transition from v3.2.1 ended with its retirement on March 31, 2024, while v4.0.1 was published in 2026 to address stakeholder feedback, as documented in the PCI Security Standards Council’s PCI DSS timeline.

The operating baseline has changed
The previously future-dated v4.x requirements became mandatory after March 31, 2025, so testing now reaches further into operational behavior. Teams need current evidence for recurring vulnerability scans, periodic penetration testing, significant-change testing, automated log review, and stronger authentication controls. The exact applicability depends on the organization’s scope and validation method, but the planning assumption should be continuous evidence production rather than a single annual event.
Verizon’s analysis makes the operational weakness clear. Its 2020 state-of-compliance review examined 77,504 PCI DSS controls across 328 validation reports, with the 2019 to 2020 comparison covering 146,496 controls in total. Requirement 11, the testing-heavy area, was fully met by the average organization globally only 60% of the time, even though its control gap improved from 13.2% to 7.4% year over year. The same analysis reported compensating controls in use for 5.8% of organizations, which means a failed primary control often creates an evidence and retest burden rather than an immediate clean resolution. See the Verizon methodology and state-of-compliance analysis for the underlying figures.
Practical rule: If a control can change through deployment, configuration, access administration, or cloud automation, its evidence needs an owner and a refresh trigger.
Payment businesses also need to understand the environment around the CDE, not only the infrastructure that stores account data. Teams responsible for card acceptance for retail shops often depend on payment providers, hosted pages, terminals, network services, and support workflows. Each dependency deserves a documented decision about whether it is in scope, connected to the CDE, or capable of affecting its security.
Defining Scope and Proving Segmentation Actually Works
Segmentation reduces testing scope only when the organization can prove that it blocks relevant paths. A network diagram that shows separate zones isn’t evidence by itself. The assessor wants to understand what data flows through the environment, which systems can reach the CDE, what identities can cross the boundary, and how the organization tested those controls.
Start with discovery, not the firewall
Build the inventory from payment flows and operational reality. Trace authorization, capture, settlement, refunds, chargebacks, support access, backups, logging, monitoring, deployment, and administrative paths. Include systems that store, process, or transmit account data, as well as components connected to the CDE or capable of affecting its security.
Use several evidence sources rather than trusting one inventory:
- Cloud inventories: Review accounts, subscriptions, projects, virtual networks, containers, orchestration, identity services, and managed databases.
- Application dependencies: Identify payment pages, APIs, reverse proxies, service meshes, webhooks, queues, third-party components, and software-as-a-service tools.
- Operational access: Map administrator workstations, jump hosts, remote support, privileged identities, build systems, repositories, and deployment pipelines.
- Data handling: Record backups, exports, test datasets, removable media, and disposal processes. Organizations should also document IT team data destruction workflows so retention and destruction decisions are testable.
Update network and data-flow diagrams from those findings. A useful diagram shows direction, protocol purpose, trust boundary, authentication dependency, data type, owner, and logging point. Avoid diagrams that only display boxes and arrows without explaining why each connection exists.

Test every meaningful boundary
Segmentation testing should begin from the perspectives an attacker could realistically obtain. That can include a corporate workstation, shared-services network, cloud workload, compromised application identity, build pipeline, or third-party access path. Network controls matter, but authentication and authorization boundaries matter too. A stolen credential that grants direct access to payment systems can defeat a diagrammed network boundary without any firewall exploit.
For each boundary, retain the source location, destination attempted, control expected to block the path, result observed, and supporting logs. Test both permitted business paths and denied paths. The denied-path evidence is especially valuable because it demonstrates that the boundary works under an actual request, not merely in configuration.
Service providers face a recurring segmentation obligation. The provided testing guidance states that segmentation testing for service providers is required every 6 months, while internal and external penetration testing must occur at least every 12 months and after any significant change. These intervals should be built into calendars and change-management workflows, not remembered during assessment preparation.
A strong result doesn’t mean every port is closed. It means the tester can explain whether an exposed service creates an exploitable route into the CDE and whether the organization can prove the control’s effectiveness. In cloud and ephemeral environments, preserve the deployment identifier, policy version, test source, target, and timestamp so the result remains interpretable after resources disappear.
Mapping PCI DSS Requirements to Testable Evidence
A requirement becomes manageable when the team converts it into a test case with a defined result and an evidence owner. Don’t create a folder called “PCI evidence” and fill it with screenshots. Create a traceable matrix that tells a reviewer what was tested, against which environment, using which method, and what happened.
For each requirement or subrequirement, record:
- The control objective and in-scope assets.
- The test procedure and expected pass condition.
- The evidence artifact, owner, collection date, and environment identifier.
- The exception, compensating control, remediation ticket, and retest result.
Requirement to Test Case and Evidence Matrix
| Requirement Group | Test Method | Evidence to Retain |
|---|---|---|
| Network security controls | Review rule sets, validate approved flows, test denied paths | Diagram, rule export, approval record, boundary-test results |
| Secure configurations | Compare hardened baselines with live systems | Configuration snapshots, deviation records, remediation tickets |
| Stored account data protection | Inspect storage locations, retention settings, and access paths | Data inventory, encryption settings, access review |
| Transmission protection | Validate approved encrypted connections and certificate configuration | Connection test output, configuration record, exception evidence |
| Malware protection | Review deployment and alert behavior where applicable | Policy, console evidence, alert test, incident record |
| Secure development | Test code review, dependency controls, deployment approval, and pipeline access | Pull-request evidence, pipeline configuration, release record |
| Access control | Review roles and attempt unauthorized access paths | Access matrix, identity logs, test results, removal evidence |
| Identification and authentication | Test authentication policy, privileged access, and account lifecycle | Policy snapshot, authentication logs, joiner-mover-leaver records |
| Physical security | Verify restricted access and media handling | Access logs, visitor records, destruction certificates |
| Logging and monitoring | Generate test events, confirm collection, alerting, protection, and review | Event samples, alert tickets, review records, retention settings |
| Security testing | Run scans, penetration tests, segmentation checks, and retests | Approved reports, findings register, remediation and retest evidence |
| Security policy and scope governance | Confirm ownership, review cadence, scope validation, and third-party responsibilities | Signed scope confirmation, policies, meeting records, responsibility matrix |
Requirement 11 deserves special handling because it is where the program demonstrates that scanning and testing operate. Verizon’s analysis found that the average organization fully complied with Requirement 11 only 60% of the time, so a passing report from one cycle shouldn’t be treated as durable proof. Retain scan coverage, failed findings, corrected results, penetration-test narratives, segmentation outcomes, and change-triggered tests together.
For safe test-data workflows, teams can also consult GoReplay’s guidance on masking production data for testing. The important control is not the product name. It’s the demonstrable process that removes or transforms sensitive values before data reaches a test environment, verifies the transformation, and prevents test systems from sending real payment transactions.
Running Vulnerability Scans Pen Tests and Safe Traffic Replay
Scans, penetration tests, and traffic replay answer different questions. A vulnerability scan asks whether known weaknesses or configuration problems are present. A penetration test asks whether an attacker can exploit weaknesses and move toward protected assets. Traffic replay asks whether the application and its control paths behave correctly under realistic requests after a change.
Build the scan schedule around coverage
External scans should cover the approved perimeter, while internal scans should cover the systems and network segments defined by scope. Review every finding, distinguish false positives with evidence, assign remediation ownership, and preserve the approved report alongside the corrected scan. A scan that runs successfully but excludes a newly deployed payment endpoint is not useful evidence.
Treat significant change as a testing trigger. Network redesigns, cloud migrations, new payment applications, authentication changes, reverse-proxy changes, and segmentation-policy updates can alter exposure even when the production deployment appears routine.
Make penetration testing adversarial
Internal and external penetration tests must be performed at least every 12 months and after any significant change, according to the provided PCI testing guidance. The test should cover network-layer and application-layer exposure, segmentation boundaries, authentication and authorization paths, privilege escalation, lateral movement, and realistic routes toward account data.
A report listing vulnerabilities without attack narratives is weak evidence. The deliverable should identify the tested perspective, target, attack path, control that failed or held, business consequence, severity rationale, remediation recommendation, and retest status. The QSA determines whether the scope and methodology satisfy the assessment, so align those expectations before testing begins.
Replay traffic without replaying sensitive data
Safe production-traffic replay is useful when it preserves application behavior while removing payment data and side effects. The pattern I recommend is:
- Capture narrowly: Collect only the endpoints, services, and request classes required to validate the change. Don’t create an uncontrolled archive of production traffic.
- Mask before storage: Transform PANs, authentication tokens, session identifiers, personal data, and other sensitive fields during capture or before persistent storage. Confirm that masking applies to headers, query strings, cookies, nested JSON, and form fields.
- Isolate the candidate: Replay into a separate environment with separate databases, credentials, queues, storage, and secrets. Disable payment submission, outbound settlement, emails, webhooks, and other irreversible actions.
- Preserve realistic relationships: Keep request sequencing and session behavior where the application requires them, but replace identities and tokens with synthetic equivalents.
- Compare safely: Evaluate status codes, response structure, validation errors, authorization outcomes, payload shape, and latency. Don’t compare raw sensitive values.
- Test boundary behavior: Send masked requests from representative non-CDE paths and observe whether reverse proxies, APIs, service meshes, and authorization layers enforce the intended restrictions.
GoReplay is one option for this workflow because it can capture and replay HTTP traffic, while its middleware supports rewriting headers, URL paths, and selected JSON fields. Teams evaluating the approach can review GoReplay’s production-traffic replay patterns and then validate the design with their QSA and data-protection owners.
The replay environment must not become a shadow CDE. Restrict access, encrypt stored captures, define retention, log administrative activity, and delete artifacts when the test ends. Replay doesn’t replace a required penetration test or ASV scan. It supplements them by exposing regressions in application-layer behavior and segmentation paths that automated checks may not exercise.
Verifying Logging Monitoring and Building Audit Ready Evidence
A control that blocks an attack but leaves no reliable record is difficult to defend during an assessment. Logging and monitoring evidence should show what happened, when it happened, which identity or system initiated it, where the event was collected, and how the organization responded.

Test the monitoring path end to end
Don’t stop at a screenshot of a logging configuration. Generate controlled events and follow them through collection, transport, storage, alerting, triage, and retention. Useful test events include an unsuccessful privileged login, a change to a security rule, an access-denied event at a segmentation boundary, a modification to a monitored file, and an administrative action in a cloud control plane.
For each event, record:
- Generation: The test action, system, identity, and timestamp.
- Collection: The source, collector, transport protection, and arrival time.
- Protection: Access restrictions, integrity safeguards, and retention configuration.
- Detection: Alert rule, routing destination, severity, and assigned responder.
- Response: Ticket, investigation notes, escalation decision, and closure evidence.
The v4.0.1 operating model increases the importance of automated log review and stronger authentication controls. That means teams should test not only whether logs exist, but whether automated review detects relevant patterns and whether privileged authentication produces enough evidence for investigation.
Keep evidence tied to environment state
An audit package should be reproducible. Retain approved scope diagrams, asset inventories, scan reports, penetration-test reports, remediation records, retest results, segmentation evidence, log samples, access reviews, and configuration snapshots. Add version or deployment identifiers for cloud policies, application builds, infrastructure-as-code, and payment-page components.
Evidence collection works best as part of change management. When a payment page changes, the ticket should identify affected controls, required tests, evidence owners, and the decision to release. When a cloud migration changes the CDE boundary, the team should refresh scope documentation and rerun relevant segmentation and application checks before relying on old evidence.
Evidence standard: A reviewer should be able to connect the control, the tested asset, the observed result, the remediation decision, and the current state without asking your team to reconstruct the story from memory.
Remediation Retesting and Working Effectively With Your QSA
Testing creates value only when findings close in a defensible loop. For every failure, create a ticket with the affected asset, control, technical evidence, risk rationale, owner, target date, compensating-control decision, and retest requirement. Prioritize exploitable paths to the CDE and weaknesses that undermine segmentation, authentication, logging, or broad control coverage.
The remediation burden can be substantial. In one large commercial penetration-testing dataset, 48% of all findings were resolved and 69% of serious findings were resolved, with a median time to resolve all severities of 67 days, as summarized by PCI DSS testing statistics and remediation guidance. These figures aren’t a forecast for every organization, but they are a useful warning against scheduling testing so close to the assessment that there’s no time for engineering work and follow-up validation.
Retest the fix, not the paperwork
A retest should reproduce the original attack or control check, confirm that the vulnerability no longer works, and look for regressions introduced by the fix. Preserve the original evidence, the remediation change, the retest method, and the final result. If a compensating control is used, document why the primary control isn’t in place, how the alternative reduces risk, who approved it, and how often its effectiveness will be checked.
Sampling decisions also need to be explicit in v4.x. Document the population, selection method, rationale, exclusions, and evidence that the sample represents the tested control. A small sample without a defensible population statement looks like convenience testing.
Bring the QSA into planning before the test starts. Confirm scope assumptions, segmentation perspectives, acceptable testing windows, evidence formats, sampling treatment, and the definition of significant change. QSAs don’t need a polished surprise at the end. They need a clear, traceable record that explains what was tested and why the result supports the compliance conclusion.
The sustainable cadence is straightforward: a failed control produces a ticket, the ticket produces a fix, the fix produces a retest, and the retest produces retained evidence. That process turns PCI DSS compliance testing into a known-good baseline that can withstand infrastructure change instead of an annual scramble.
GoReplay can help teams capture and replay masked HTTP traffic in isolated test environments, giving segmentation and application teams a practical way to validate behavior beyond routine scans. Visit GoReplay to evaluate how its capture, masking, and replay workflows can fit into your PCI DSS evidence and retesting program.