Web Application Penetration Testing
A phased pentesting workflow for running web applications. Adapted from
Shannon's pipeline (Keygraph, AGPL — concepts only, no code borrowed).
Built around three rules:
- No exploit, no report — every finding requires reproducible evidence.
- Bounded scope — every active request goes against a target the operator
pre-declared. Off-scope hosts are refused.
- Bypass exhaustion before false-positive dismissal — a "blocked" payload
is not a clean bill of health until you've tried the bypass set.
⚠️ Hard Guardrails — Read Before Every Engagement
Violating any of these invalidates the engagement and may be illegal.
Authorization gate. Before the first active scan in a session, you
MUST confirm with the user, in writing, that they own or have written
authorization to test the target. Record the acknowledgement in
engagement/authorization.md (see template). No acknowledgement → no
active scanning. Reading public pages with curl is fine; sending
payloads is not.
Scope allowlist. Maintain engagement/scope.txt — one hostname or
CIDR per line. Every nmap, curl, whatweb, browser navigation, or
payload-bearing request MUST be against an entry in scope. If a target
redirects you off-scope (3xx to a different host, a link in HTML),
STOP and confirm with the user before following.
No production systems without paper. If the user hasn't told you
"yes, prod is in scope and I have written sign-off," assume not. Default
targets are staging, local docker, dedicated test instances.
Cloud metadata is off by default. Do not probe 169.254.169.254,
metadata.google.internal, 100.100.100.200, [fd00:ec2::254], or
equivalent unless the engagement explicitly includes SSRF-to-metadata
as a goal AND the target is one you control. The agent's browser tool
can reach these from inside your own infrastructure — don't.
Destructive payloads need approval. SQLi payloads that DROP/DELETE,
filesystem-write SSTI, command injection with rm/shutdown/mkfs,
anything that mutates beyond a single test row → ASK FIRST. The
approval.py system catches some; don't rely on it alone.
Aux-client leakage risk (Hermes-specific). This skill produces
sessions full of SQLi/XSS/RCE payloads, captured credentials, JWT
tokens. Hermes' compression and title-generation paths replay history
through the auxiliary client (often the main model). Anything sensitive
you write to the conversation can leave the box on the next compress.
Mitigation:
- Redact captured tokens/credentials to the LAST 6 CHARS before logging
them in any message. Full values go to
engagement/evidence/ files,
never into chat history.
- If the engagement is sensitive, set
auxiliary.title_generation.enabled: false
in ~/.hermes/config.yaml for the session.
Rate limit yourself. Default 200ms between active requests against
any single host. The recon-scan.sh script enforces this. Don't bypass
it without operator approval.
Authority of the report. This skill produces a security
assessment, not a "PASS." Even a clean run is "no exploitable issues
FOUND in scope X within time T using methods Y" — not "the application
is secure." Mirror that language in the report.
Phase 0: Engagement Setup
Before any scanning happens, create the engagement directory and
authorization acknowledgement.
ENGAGEMENT=engagement-$(date +%Y%m%d-%H%M%S)
mkdir -p "$ENGAGEMENT"/{evidence,findings,reports}
cd "$ENGAGEMENT"
Ask the user (verbatim):
"Confirm: (a) the target URL is [X], (b) you own this application
or have written authorization to test it, and (c) the engagement
may run for up to [N] hours starting now. Reply 'authorized' to
proceed."
Wait for explicit authorized response. Any other answer means STOP.
Record authorization to engagement/authorization.md using the
template in templates/authorization.md. Include:
- Target URL(s) and IP(s)
- Authorization basis (ownership / written authz from $name)
- Engagement window
- Out-of-scope items (production, third-party services, etc.)
- Operator name (the user driving this session)
Build scope.txt:
localhost
127.0.0.1
staging.example.com
192.168.1.0/24 # internal lab only, with operator OK
Read references/scope-enforcement.md before issuing the first
active request — that doc has the host-extraction rules you apply
to every command/URL before it goes out.
Phase 1: Pre-Recon (Code Analysis, optional)
Skip if no source access (black-box engagement).
If you have read access to the application source:
- Map the architecture — framework, routing, middleware stack
- Inventory sinks — every
execute(, os.system(, eval(,
template render, file read/write, redirect target
- Map auth — session cookie vs JWT, OAuth flows, password reset,
privileged endpoints
- Identify trust boundaries — what's authenticated, what's not,
what comes from
request.*
- Backward taint from each sink to a request source. Early-terminate
when proper sanitization is found (parameterized queries, allowlists,
shlex.quote, well-known escapers).
Output: evidence/pre-recon.md — architecture map, sink inventory,
suspected vulnerable code paths.
This is OFFLINE work. No traffic to the target.
Phase 2: Recon (Live, Read-Only)
Maps the attack surface. All requests are GETs of public pages, no
payloads yet. Still scope-bounded.
Verify scope. Resolve every target hostname → IP. Confirm IPs are
in scope (avoids the "DNS points somewhere unexpected" trap).
Network surface (only if scope permits port scanning):
nmap -sT -T3 --top-ports 100 -oN evidence/nmap.txt $TARGET
Use -T3 (default), not -T4/-T5. Stealthier and avoids tripping
IDS/IPS in shared environments.
Tech fingerprint:
whatweb -v $TARGET_URL > evidence/whatweb.txt
curl -sIk $TARGET_URL > evidence/headers.txt
Endpoint discovery:
- Crawl the app with the browser tool (
browser_navigate,
browser_get_images, follow links).
- Inspect
robots.txt, sitemap.xml, .well-known/*.
- Use the developer tools network panel via browser tool to capture
XHR/fetch calls.
Auth surface: Identify login, registration, password reset,
session cookie names, token formats. Do NOT send credentials yet —
just observe.
Correlate with pre-recon (if you have source). For each
evidence/pre-recon.md finding, mark whether the live surface
confirms it's reachable.
Output: evidence/recon.md — endpoints, technologies, auth model,
input vectors.
Phase 3: Vulnerability Analysis
One delegate_task per vulnerability class. Each agent reads
evidence/recon.md (+ evidence/pre-recon.md if present), produces
findings/<class>-queue.json using templates/exploitation-queue.json.
Use delegate_task with these focused subagents (parallel where possible):
| Class |
Goal |
Reference |
injection |
SQLi, command, path traversal, SSTI, LFI/RFI, deserialization |
references/vuln-taxonomy.md (slot types) |
xss |
Reflected, stored, DOM-based |
references/vuln-taxonomy.md (render contexts) |
auth |
Login bypass, JWT confusion, session fixation, OAuth flaws |
references/exploitation-techniques.md |
authz |
IDOR, vertical/horizontal escalation, business logic |
references/exploitation-techniques.md |
ssrf |
Internal reachability, metadata, protocol smuggling |
Skip metadata unless explicitly authorized |
infra |
Misconfig, info disclosure, default creds, exposed admin |
references/exploitation-techniques.md |
Each queue entry has: id, vuln class, source (file:line if known),
endpoint, parameter, slot type, suspected defense, verdict
(identified / partial / confirmed / critical), witness payload,
confidence (0-1), notes.
The analysis phase doesn't send malicious payloads yet — it stages them.
The exploitation phase actually fires them.
Phase 4: Exploitation (Proof-Based, Conditional)
Only run a sub-agent per class where the analysis queue has actionable
entries (identified or partial).
For each candidate:
- Pre-send check — host in scope? auth gate satisfied? payload
approved if destructive?
- Send the witness payload — minimal proof. SQLi:
' AND 1=1--
then ' AND 1=2--. XSS: a benign marker like
<svg/onload=console.log("HERMES-PENTEST-XSS")>. Never alert(1) in
stored XSS — it'll fire for other users in shared environments.
- Verify the witness fires — for blind injection, use a sleep
probe (
SLEEP(5)) and time the response. For SSRF, use a
tester-controlled callback host you own (NOT a public service like
webhook.site for sensitive engagements — exfil paths).
- Promote level:
- L1 Identified — pattern matched, no behavior change
- L2 Partial — sink reached, but defense in place
- L3 Confirmed — payload changed app behavior in observable way
- L4 Critical — data extracted, code executed, access escalated
- Bypass exhaustion before classifying as FP. For each candidate
that blocks: try at least the bypass set in
references/bypass-techniques.md for that class. Only after the set
is exhausted may you write verdict: false_positive.
- Record evidence for every L3/L4:
- Full request (method, URL, headers, body)
- Response (status, headers, relevant body excerpt)
- Reproducer command (curl one-liner)
- Impact statement
Output: findings/exploitation-evidence.md
Redact in evidence files:
- Any captured credentials/tokens → last 6 chars only in chat;
full value to
findings/secrets-vault.md (gitignored).
- Other users' PII → redact.
- Your test credentials → fine to keep.
Phase 5: Reporting
Generate the final report using templates/pentest-report.md. Sections:
- Executive summary
- Engagement scope (from
engagement/scope.txt)
- Authorization (from
engagement/authorization.md)
- Findings (L3/L4 only — proof-required). Per finding:
- Title, severity (CVSS 3.1), CWE
- Affected endpoint(s)
- Proof (request + response excerpt)
- Reproduction steps
- Impact
- Remediation
- Not-exploited candidates (L1/L2 with notes on what blocked them)
- Out-of-scope observations
- Methodology / tools used
- Limitations and what was NOT tested
Severity policy: CVSS only for L3/L4. L1/L2 are "candidates pending
verification" — don't assign CVSS to unverified findings.
When to Stop
- The user revokes authorization.
- A candidate finding clearly impacts production data and you don't have
approval for destructive testing — STOP and ask.
- The target starts returning 503/429 storms — back off, reconvene with
the operator.
- You discover something outside the contracted scope (e.g. an exposed
customer database while testing an unrelated endpoint). STOP, document,
report to the operator. Do not pivot without explicit approval — that
pivot is what makes pentesting illegal.
What This Skill Does NOT Cover
- Network-layer pentesting beyond port scanning (no Metasploit,
Cobalt Strike, AD attacks, network protocol fuzzing).
- Reverse engineering / binary analysis (see issue #383).
- Source-only static analysis (see issue #382).
- Active social engineering / phishing.
- Anything against systems the operator hasn't pre-authorized.
If the engagement needs any of these, escalate to a professional
pentester. This skill complements professional pentesting; it does
not replace it.
Further Reading
references/scope-enforcement.md — how to bound every active request
references/vuln-taxonomy.md — slot types, render contexts, OWASP map
references/exploitation-techniques.md — per-class payload patterns
references/bypass-techniques.md — common WAF/filter bypasses
templates/authorization.md — engagement authorization template
templates/pentest-report.md — final report template
templates/exploitation-queue.json — per-class finding queue schema
scripts/recon-scan.sh — rate-limited nmap+whatweb+headers wrapper
Source: NousResearch/hermes-agent → optional-skills/security/web-pentest/SKILL.md
1---2name: web-pentest3description: | Authorized web application penetration testing — reconnaissance, vulnerability analysis, proof-based exploitation, and professional reporting. Adapts Shannon's "No Exploit, No Report" methodology with hard guardrails for scope, authorization, and aux-client leakage. Active testing against running applications you own or have written authorization to test.4---567# Web Application Penetration Testing89A phased pentesting workflow for running web applications. Adapted from10Shannon's pipeline (Keygraph, AGPL — concepts only, no code borrowed).11Built around three rules:12131. No exploit, no report — every finding requires reproducible evidence.142. Bounded scope — every active request goes against a target the operator15 pre-declared. Off-scope hosts are refused.163. Bypass exhaustion before false-positive dismissal — a "blocked" payload17 is not a clean bill of health until you've tried the bypass set.1819---2021## ⚠️ Hard Guardrails — Read Before Every Engagement2223Violating any of these invalidates the engagement and may be illegal.24251. **Authorization gate.** Before the first active scan in a session, you26 MUST confirm with the user, in writing, that they own or have written27 authorization to test the target. Record the acknowledgement in28 `engagement/authorization.md` (see template). No acknowledgement → no29 active scanning. Reading public pages with `curl` is fine; sending30 payloads is not.31322. **Scope allowlist.** Maintain `engagement/scope.txt` — one hostname or33 CIDR per line. Every `nmap`, `curl`, `whatweb`, browser navigation, or34 payload-bearing request MUST be against an entry in scope. If a target35 redirects you off-scope (3xx to a different host, a link in HTML),36 STOP and confirm with the user before following.37383. **No production systems without paper.** If the user hasn't told you39 "yes, prod is in scope and I have written sign-off," assume not. Default40 targets are staging, local docker, dedicated test instances.41424. **Cloud metadata is off by default.** Do not probe `169.254.169.254`,43 `metadata.google.internal`, `100.100.100.200`, `[fd00:ec2::254]`, or44 equivalent unless the engagement explicitly includes SSRF-to-metadata45 as a goal AND the target is one you control. The agent's browser tool46 can reach these from inside your own infrastructure — don't.47485. **Destructive payloads need approval.** SQLi payloads that DROP/DELETE,49 filesystem-write SSTI, command injection with `rm`/`shutdown`/`mkfs`,50 anything that mutates beyond a single test row → ASK FIRST. The51 `approval.py` system catches some; don't rely on it alone.52536. **Aux-client leakage risk (Hermes-specific).** This skill produces54 sessions full of SQLi/XSS/RCE payloads, captured credentials, JWT55 tokens. Hermes' compression and title-generation paths replay history56 through the auxiliary client (often the main model). Anything sensitive57 you write to the conversation can leave the box on the next compress.58 Mitigation:59 - Redact captured tokens/credentials to the LAST 6 CHARS before logging60 them in any message. Full values go to `engagement/evidence/` files,61 never into chat history.62 - If the engagement is sensitive, set `auxiliary.title_generation.enabled: false`63 in `~/.hermes/config.yaml` for the session.64657. **Rate limit yourself.** Default 200ms between active requests against66 any single host. The recon-scan.sh script enforces this. Don't bypass67 it without operator approval.68698. **Authority of the report.** This skill produces a security70 assessment, not a "PASS." Even a clean run is "no exploitable issues71 FOUND in scope X within time T using methods Y" — not "the application72 is secure." Mirror that language in the report.7374---7576## Phase 0: Engagement Setup7778Before any scanning happens, create the engagement directory and79authorization acknowledgement.8081```bash82ENGAGEMENT=engagement-$(date +%Y%m%d-%H%M%S)83mkdir -p "$ENGAGEMENT"/{evidence,findings,reports}84cd "$ENGAGEMENT"85```86871. **Ask the user (verbatim):**88 > "Confirm: (a) the target URL is [X], (b) you own this application89 > or have written authorization to test it, and (c) the engagement90 > may run for up to [N] hours starting now. Reply 'authorized' to91 > proceed."92932. **Wait for explicit `authorized` response.** Any other answer means STOP.94953. **Record authorization** to `engagement/authorization.md` using the96 template in `templates/authorization.md`. Include:97 - Target URL(s) and IP(s)98 - Authorization basis (ownership / written authz from $name)99 - Engagement window100 - Out-of-scope items (production, third-party services, etc.)101 - Operator name (the user driving this session)1021034. **Build scope.txt:**104 ```105 localhost106 127.0.0.1107 staging.example.com108 192.168.1.0/24 # internal lab only, with operator OK109 ```1101115. **Read** `references/scope-enforcement.md` before issuing the first112 active request — that doc has the host-extraction rules you apply113 to every command/URL before it goes out.114115---116117## Phase 1: Pre-Recon (Code Analysis, optional)118119Skip if no source access (black-box engagement).120121If you have read access to the application source:1221231. **Map the architecture** — framework, routing, middleware stack1242. **Inventory sinks** — every `execute(`, `os.system(`, `eval(`,125 template render, file read/write, redirect target1263. **Map auth** — session cookie vs JWT, OAuth flows, password reset,127 privileged endpoints1284. **Identify trust boundaries** — what's authenticated, what's not,129 what comes from `request.*`1305. **Backward taint** from each sink to a request source. Early-terminate131 when proper sanitization is found (parameterized queries, allowlists,132 `shlex.quote`, well-known escapers).133134Output: `evidence/pre-recon.md` — architecture map, sink inventory,135suspected vulnerable code paths.136137This is OFFLINE work. No traffic to the target.138139---140141## Phase 2: Recon (Live, Read-Only)142143Maps the attack surface. All requests are GETs of public pages, no144payloads yet. Still scope-bounded.1451461. **Verify scope.** Resolve every target hostname → IP. Confirm IPs are147 in scope (avoids the "DNS points somewhere unexpected" trap).1481492. **Network surface** (only if scope permits port scanning):150 ```bash151 nmap -sT -T3 --top-ports 100 -oN evidence/nmap.txt $TARGET152 ```153 Use `-T3` (default), not `-T4/-T5`. Stealthier and avoids tripping154 IDS/IPS in shared environments.1551563. **Tech fingerprint:**157 ```bash158 whatweb -v $TARGET_URL > evidence/whatweb.txt159 curl -sIk $TARGET_URL > evidence/headers.txt160 ```1611624. **Endpoint discovery:**163 - Crawl the app with the browser tool (`browser_navigate`,164 `browser_get_images`, follow links).165 - Inspect `robots.txt`, `sitemap.xml`, `.well-known/*`.166 - Use the developer tools network panel via browser tool to capture167 XHR/fetch calls.1681695. **Auth surface:** Identify login, registration, password reset,170 session cookie names, token formats. Do NOT send credentials yet —171 just observe.1721736. **Correlate with pre-recon** (if you have source). For each174 `evidence/pre-recon.md` finding, mark whether the live surface175 confirms it's reachable.176177Output: `evidence/recon.md` — endpoints, technologies, auth model,178input vectors.179180---181182## Phase 3: Vulnerability Analysis183184One delegate_task per vulnerability class. Each agent reads185`evidence/recon.md` (+ `evidence/pre-recon.md` if present), produces186`findings/<class>-queue.json` using `templates/exploitation-queue.json`.187188Use `delegate_task` with these focused subagents (parallel where possible):189190| Class | Goal | Reference |191|-------|------|-----------|192| `injection` | SQLi, command, path traversal, SSTI, LFI/RFI, deserialization | `references/vuln-taxonomy.md` (slot types) |193| `xss` | Reflected, stored, DOM-based | `references/vuln-taxonomy.md` (render contexts) |194| `auth` | Login bypass, JWT confusion, session fixation, OAuth flaws | `references/exploitation-techniques.md` |195| `authz` | IDOR, vertical/horizontal escalation, business logic | `references/exploitation-techniques.md` |196| `ssrf` | Internal reachability, metadata, protocol smuggling | Skip metadata unless explicitly authorized |197| `infra` | Misconfig, info disclosure, default creds, exposed admin | `references/exploitation-techniques.md` |198199Each queue entry has: id, vuln class, source (file:line if known),200endpoint, parameter, slot type, suspected defense, verdict201(`identified` / `partial` / `confirmed` / `critical`), witness payload,202confidence (0-1), notes.203204The analysis phase doesn't send malicious payloads yet — it stages them.205The exploitation phase actually fires them.206207---208209## Phase 4: Exploitation (Proof-Based, Conditional)210211Only run a sub-agent per class where the analysis queue has actionable212entries (`identified` or `partial`).213214For each candidate:2152161. **Pre-send check** — host in scope? auth gate satisfied? payload217 approved if destructive?2182. **Send the witness payload** — minimal proof. SQLi: `' AND 1=1--`219 then `' AND 1=2--`. XSS: a benign marker like220 `<svg/onload=console.log("HERMES-PENTEST-XSS")>`. Never `alert(1)` in221 stored XSS — it'll fire for other users in shared environments.2223. **Verify the witness fires** — for blind injection, use a sleep223 probe (`SLEEP(5)`) and time the response. For SSRF, use a224 tester-controlled callback host you own (NOT a public service like225 webhook.site for sensitive engagements — exfil paths).2264. **Promote level:**227 - **L1 Identified** — pattern matched, no behavior change228 - **L2 Partial** — sink reached, but defense in place229 - **L3 Confirmed** — payload changed app behavior in observable way230 - **L4 Critical** — data extracted, code executed, access escalated2315. **Bypass exhaustion before classifying as FP.** For each candidate232 that blocks: try at least the bypass set in233 `references/bypass-techniques.md` for that class. Only after the set234 is exhausted may you write `verdict: false_positive`.2356. **Record evidence** for every L3/L4:236 - Full request (method, URL, headers, body)237 - Response (status, headers, relevant body excerpt)238 - Reproducer command (curl one-liner)239 - Impact statement240241Output: `findings/exploitation-evidence.md`242243**Redact in evidence files:**244- Any captured credentials/tokens → last 6 chars only in chat;245 full value to `findings/secrets-vault.md` (gitignored).246- Other users' PII → redact.247- Your test credentials → fine to keep.248249---250251## Phase 5: Reporting252253Generate the final report using `templates/pentest-report.md`. Sections:2542551. Executive summary2562. Engagement scope (from `engagement/scope.txt`)2573. Authorization (from `engagement/authorization.md`)2584. Findings (L3/L4 only — proof-required). Per finding:259 - Title, severity (CVSS 3.1), CWE260 - Affected endpoint(s)261 - Proof (request + response excerpt)262 - Reproduction steps263 - Impact264 - Remediation2655. Not-exploited candidates (L1/L2 with notes on what blocked them)2666. Out-of-scope observations2677. Methodology / tools used2688. Limitations and what was NOT tested269270**Severity policy:** CVSS only for L3/L4. L1/L2 are "candidates pending271verification" — don't assign CVSS to unverified findings.272273---274275## When to Stop276277- The user revokes authorization.278- A candidate finding clearly impacts production data and you don't have279 approval for destructive testing — STOP and ask.280- The target starts returning 503/429 storms — back off, reconvene with281 the operator.282- You discover something *outside* the contracted scope (e.g. an exposed283 customer database while testing an unrelated endpoint). STOP, document,284 report to the operator. Do not pivot without explicit approval — that285 pivot is what makes pentesting illegal.286287---288289## What This Skill Does NOT Cover290291- Network-layer pentesting beyond port scanning (no Metasploit,292 Cobalt Strike, AD attacks, network protocol fuzzing).293- Reverse engineering / binary analysis (see issue #383).294- Source-only static analysis (see issue #382).295- Active social engineering / phishing.296- Anything against systems the operator hasn't pre-authorized.297298If the engagement needs any of these, escalate to a professional299pentester. This skill complements professional pentesting; it does300not replace it.301302---303304## Further Reading305306- `references/scope-enforcement.md` — how to bound every active request307- `references/vuln-taxonomy.md` — slot types, render contexts, OWASP map308- `references/exploitation-techniques.md` — per-class payload patterns309- `references/bypass-techniques.md` — common WAF/filter bypasses310- `templates/authorization.md` — engagement authorization template311- `templates/pentest-report.md` — final report template312- `templates/exploitation-queue.json` — per-class finding queue schema313- `scripts/recon-scan.sh` — rate-limited nmap+whatweb+headers wrapper314315---316317**Source:** [`NousResearch/hermes-agent`](https://github.com/NousResearch/hermes-agent) → `optional-skills/security/web-pentest/SKILL.md`