Skip to content
Red Cell

How to Read a Penetration Testing Report

Kenneth Brown

Definition · Published · 4 min read

Two reports in one

The National Institute of Standards and Technology (NIST) describes the end of a penetration test this way in Special Publication 800-115: "At the conclusion of the test, a report is generally developed to describe identified vulnerabilities, present a risk rating, and give guidance on how to mitigate the discovered weaknesses."

The same publication adds: "Because a report may have multiple audiences, multiple report formats may be required to ensure that all are appropriately addressed." In practice, a useful report is written twice over: an executive summary for the people who decide what gets funded, and technical findings for the people who make the change. Read them in that order, and expect each to stand on its own.

If you are still choosing a tester, what a penetration test is covers what the engagement should include before you get this far.

Reading the executive summary

The executive summary should answer four questions without jargon:

  • What was tested, and over what window.
  • What an attacker could reach, stated as outcomes: customer records, administrative control, payment data, another tenant's account.
  • What that means for the business, in the terms your leadership already uses.
  • What to do first.

Be wary of a summary that is mostly a count of findings by severity. Twelve medium findings and one critical tell you very little until someone explains which of them lead anywhere.

Reading a technical finding

Each finding should be complete enough that an engineer who was not on the engagement can confirm it, fix it and check the fix.

PartWhat it should tell youWarning sign
Affected componentThe specific host, endpoint, role or configuration involvedA whole application or network named as affected
Severity and rationaleThe rating, and why, including what an attacker needs and what they gainA score with no explanation of how it applies to you
Conditions requiredAccess level, user role or network position needed to exploit itNo mention of preconditions
EvidenceRequests, responses, screenshots or configuration excerpts that prove itTool output pasted without confirmation
Reproduction stepsA sequence a developer can follow to see the issue themselvesSteps that only the tester could repeat
RemediationThe change that removes the issue, at the level of the controlGeneric advice such as apply best practice
The parts of a technical finding and what to look for in each.

Evidence and reproduction steps

Evidence is what separates a finding from an opinion. It should show the issue in your environment, with sensitive values redacted, and no more data than is needed to prove the point. Reproduction steps turn that evidence into something your team can verify for itself. If your engineers cannot reproduce a finding from the steps given, ask the tester to walk through it before the report is final.

Remediation guidance

NIST SP 800-115 says mitigation recommendations should be developed "for each finding," and adds: "There may be both technical recommendations (e.g., applying a particular patch) and nontechnical recommendations that address the organization’s processes (e.g., updating the patch management process)."

The most useful remediation addresses the cause, not the single request that demonstrated it. A finding about one endpoint that skips an authorization check usually points to an authorization model that needs to be enforced in one place. Good guidance says so, and suggests an order of work across findings.

Severity, scores and real reachability

Many reports rate findings with the Common Vulnerability Scoring System (CVSS), maintained by the Forum of Incident Response and Security Teams (FIRST). FIRST describes CVSS as "an open framework for communicating the characteristics and severity of software vulnerabilities." It has four metric groups: Base, Threat, Environmental and Supplemental. The Base group "represents the intrinsic characteristics of a vulnerability that are constant over time and across user environments."

That last phrase matters. FIRST's CVSS version 4.0 User Guide states that "CVSS Base (CVSS-B) scores are designed to measure the severity of a vulnerability and should not be used alone to assess risk." The specification asks that scores be labelled by the metrics used: CVSS-B for Base only, CVSS-BT with Threat, CVSS-BE with Environmental, and CVSS-BTE with all three. A score labelled CVSS-B says nothing about your environment.

FIRST also publishes an optional qualitative scale that maps scores to ratings:

RatingCVSS score
None0.0
Low0.1 – 3.9
Medium4.0 – 6.9
High7.0 – 8.9
Critical9.0 – 10.0
CVSS v4.0 qualitative severity rating scale, as published by FIRST.

A penetration test adds what a score cannot: whether the weakness was actually reachable, what it led to, and whether it chained with something else. A high-scoring issue on a host that nothing can reach may deserve less urgency than a medium-rated issue on your login path that opened the way to an administrative account. The severity rationale in each finding is where the tester should make that case.

For vulnerabilities with a published Common Vulnerabilities and Exposures (CVE) identifier, the Cybersecurity and Infrastructure Security Agency (CISA) maintains the Known Exploited Vulnerabilities (KEV) Catalog, which it describes as "the authoritative source of vulnerabilities that have been exploited in the wild." CISA says organizations "should use the KEV catalog as an input to their vulnerability management prioritization framework."

Limitations

Every report should state its limits plainly: the scope, the testing window, the systems excluded, any technique that was not permitted, and anything that could not be tested and why, such as an environment that was unavailable or credentials that arrived late.

The boundaries themselves are set before testing starts, in the statement of work (SOW) and the rules of engagement. What belongs in those documents is worth reading alongside this section of the report.

The retest

After you fix the agreed findings, the tester checks them again. Retest results should say, finding by finding, whether the issue is resolved, partly resolved or still present, with fresh evidence for each.

Check the terms before you rely on them. The SOW should say which findings a retest covers, how many verification cycles are included and within what period. At Red Cell, a retest is one verification cycle for the agreed findings, within the period and scope defined in the SOW, as our approach page describes.

Questions to ask when the report arrives

Checks before you accept a penetration testing report

  • Does the executive summary say what an attacker could reach, not just how many findings there are?
  • Does every finding name the specific affected component and the conditions required?
  • Does each severity rating come with a rationale that reflects reachability in our environment?
  • Can our engineers reproduce each finding from the steps and evidence given?
  • Does the remediation address the underlying control, with a suggested order of work?
  • Does the report state its scope, window, exclusions and anything that could not be tested?
  • Are the retest terms, cycles and window clear from the statement of work?

Most of these can be checked before you hire, by asking for a sample report. How to write a penetration testing request for proposal (RFP) covers what else to ask for at that stage.

Sources

  1. 01NIST, Special Publication 800-115, Technical Guide to Information Security Testing and AssessmentAccessed
  2. 02NIST, SP 800-115 full text (sections 5.2, 8.1 and 8.2)Accessed
  3. 03FIRST, CVSS v4.0 Specification DocumentAccessed
  4. 04FIRST, CVSS v4.0 User GuideAccessed
  5. 05CISA, Known Exploited Vulnerabilities CatalogAccessed

Frequently asked questions

Should we fix findings in the order of their severity score?
Start there, but do not stop there. A lower-rated issue on an internet-facing login path, or one that chains into something worse, can matter more than a higher-rated issue on a system nobody can reach. A good report explains that reasoning and suggests an order of work.
What if we disagree with a finding or its severity?
Raise it with the tester before the report is final. Bring the context they may not have had, such as a compensating control or the fact that a system is being retired, and ask them to record the discussion in the finding rather than quietly deleting it.
Why does the report include low and informational findings?
They record weaknesses that are not directly exploitable on their own but are worth knowing about, such as verbose error messages or missing hardening. Some of them are the first link in a chain. They should not crowd out the findings that need action now.
Can we share the report with customers or auditors?
Decide during scoping who the report is for, because that shapes how it is written. Treat the full technical findings as sensitive, since they describe how to attack your systems, and share them only under confidentiality. Many requests can be met with the executive summary and the retest results.
What does a retest actually prove?
That the specific findings retested are resolved, or not, at the time of the retest. It is not a new penetration test and does not look for new issues outside the agreed findings.

Tell us what you need to test.

Send the systems, the timing and the constraints. You get a scoped proposal, not a sales sequence.