Offensive OSINT — External Red-Team Arsenal
Companion skill:
osint-methodology(the "how to think" skill). This skill is the "what to reach for." Use them together.
0. When to use / When NOT
Use this skill when:
- You need concrete probe paths, wordlists, regexes, payloads, scoring rules, or tool URLs.
- You're executing reconnaissance and need the actual technical reference (vs. methodology).
- You're building a recon automation and need specific lists to seed it.
Do NOT use this skill when:
- The user is asking for active exploitation, post-exploitation, or anything past reconnaissance.
- The user is asking for defensive / blue-team detections.
- The target's authorization isn't established — see §1.
1. Authorization & Legal Posture
For assets the operator owns or has written authorization to assess. Soft scope check before acting against an unverified third-party target — see methodology skill §1 for the full posture.
2. Confidence Levels
- TENTATIVE — plausible based on indirect evidence (snippet-only dork match, single-source asset, inferred email pattern).
- FIRM — directly observed (subdomain resolves, HEAD-confirmed bucket exists, banner returned).
- CONFIRMED — verified via independent corroboration OR direct verification (live PMAK validation, multiple sources agree, listable bucket with object retrieval).
3. Output Format Conventions
Findings should carry: id, module, asset_key, category, severity (info/low/medium/high/critical), confidence, title, description, evidence (url + UTC timestamp + sha256 + raw ≤ 2 KiB), references, remediation. UTC timestamps everywhere.
4. Source Hygiene & Citations
URL + UTC timestamp + SHA-256 + tool version + run_id, every artifact. PNG screenshots, JSONL run logs, raw HTTP captures capped at 2 KiB body.
5. Do NOT
- Don't paste creds/PII/session tokens into cloud LLMs.
- Don't run destructive probes outside DEEP/
--aggressive. - Don't use validated credentials for anything except read-only liveness check.
- Don't single-source attribute.
- Don't assume vendor labels are ground truth.
6. General OSINT (curated tool refs)
- OSINT Bookmarks — comprehensive bookmarks.
- OSINT Framework — tool/resource directory.
- IntelTechniques Tools — investigative suite.
- Bellingcat Toolkit — investigative journalism.
- CyberSudo OSINT Toolkit — OSINT websites list.
- Google Dorks — efficient Google searching.
- Distributed Denial of Secrets — leaked datasets.
- Country-Specific Resources — country-targeted OSINT.
7. Search Engines
| Tool | Notes |
|---|---|
| Carrot2 | Clusters results by topic |
| etools | Metasearch |
| Kagi | Privacy-first, non-personalized |
| Brave Search | Independent index; Goggles for custom ranking |
| PDF Search | PDF + table of contents |
| Google Fact Check Explorer | Cross-site fact-check |
8. Username & Email Investigation
| Tool | Purpose |
|---|---|
| Sherlock | Username search across social networks |
| Maigret | Profile collector by username |
| What's My Name | Username search |
| Holehe | Email registration check |
| Epieos | Email pivots and metadata |
| OSINT Industries | Email/username/phone lookups |
| Hunter.io | Domain → emails |
| EmailRep | Email reputation |
| Emailable | Email verification |
| Mugetsu | X/Twitter username history |
| RocketReach / Apollo | Email enrichment + pattern guessing |
| PhoneInfoga | Phone number intelligence |
Browser extensions: GetProspect, SignalHire.
9. People Search
- TruePeopleSearch — free U.S. people search.
- WhitePages, Spokeo, Webmii, Pipl (paid).
- Clearbit — company/individual data enrichment.
- FaceCheck / FaceSeek — reverse face search.
10. Phone Number OSINT
- TrueCaller — caller ID + spam blocking.
- ThatsThem — reverse phone search.
- Infobel — non-USA phone search.
- FreeCarrierLookup — carrier/type (US).
- NumlookupAPI [Freemium] — programmatic carrier checks.
- CallerIDTest, Advanced Background Checks.
11. Email-Pattern Inference (TENTATIVE candidates)
Given a (first_name, last_name, domain), generate these 8 candidate addresses for breach pre-hits, phishing list curation, and downstream enrichment. Mark as TENTATIVE confidence until corroborated.
{first}.{last}@{domain} # john.doe@example.com
{first}{last}@{domain} # johndoe@example.com
{first}@{domain} # john@example.com
{first[0]}{last}@{domain} # jdoe@example.com
{first}.{last[0]}@{domain} # john.d@example.com
{last}@{domain} # doe@example.com
{first}_{last}@{domain} # john_doe@example.com
{first}-{last}@{domain} # john-doe@example.com
Lowercase before lookup. Strip diacritics for ASCII fallback. If the org uses a known pattern (e.g., Hunter.io shows {first}.{last} is dominant), prioritize that one and mark FIRM.
12. Email-Harvest Source Stack
Six parallel sources, dedup at the end:
- IntelX phonebook API — 2-step search + poll. Largest single source for breach-era addresses.
- Hunter.io — domain-search endpoint. ~25 free/month. Returns verified emails + roles.
- crt.sh — extract X.509 SAN extensions. Many certs include admin/contact emails.
- DuckDuckGo SERP scrape — HTML scrape of
"@{target-domain}"results. - Bing SERP scrape — same query, complementary index.
- Wayback CDX — historic snapshots of the target's homepage / contact / about pages often contain emails removed from the live site.
Email regex:
\b[A-Za-z0-9._%+\-]+@[A-Za-z0-9.\-]+\.[A-Za-z]{2,}\b
Noise filter (reject numeric-only locals):
^[0-9]+$
(Discards garbage like 12345@example.com from random tokens.)
13. Social Media
| Platform | Tool |
|---|---|
| Picuki — profile view without account | |
| X/Twitter | snscrape — preferred CLI scraper; Twint as fallback |
| Graph Search, sowsearch.info, lookup-id.com, whopostedwhat.com | |
| Facebook (research) | Meta Content Library — CrowdTangle successor (researcher-gated) |
| YouTube/Twitch | Social Blade — analytics |
| TikTok | Tokboard — trends + profile analytics |
| Reveddit — removed content; RedTrack.social — user history | |
| Bluesky | Firesky — real-time firehose; SkyView — follower graphs |
| Mastodon | FediSearch — cross-instance search; Fedifinder — find Twitter users on Mastodon |
| Faces | Search4Faces |
14. Public Records & Company Information
- OpenCorporates — world's largest open company DB.
- SEC EDGAR — U.S. company filings.
- OpenOwnership Register — beneficial ownership.
- MuckRock — FOIA repository + request tracking.
- EU Tenders (TED) — EU procurement notices.
- World Bank Projects — project + procurement records.
- UK Companies House — UK companies + officers + filings.
14.1 RU registries
Rusprofile, Kontur.Focus (freemium), zakupki.gov.ru (procurement), EGRUL/EGRIP (official, captcha-gated).
14.2 CN registries + USCC + ICP
- GSXT — gsxt.gov.cn National Enterprise Credit Info; cross-check with Tianyancha / Qichacha.
- USCC (Unified Social Credit Code) — 18-character entity ID assigned to all CN legal entities. Format:
<region:6><authority:2><type:1><serial:9>. Useful for joining GSXT records to ICP filings. - ICP Beian — beian.miit.gov.cn — every domain serving traffic in mainland CN must register an ICP filing; the filing links the domain to a USCC, which links to the legal entity in GSXT.
- Workflow:
target.cndomain → ICP lookup → USCC → GSXT → entity name + officers + adjacent registered entities.
14.3 Sanctions & Compliance
- OFAC SDN List, EU Sanctions Map.
- OpenSanctions — aggregated.
- OCCRP Aleph — investigative documents, leaks, company records.
15. Breach & Leak Data
- Have I Been Pwned — breach lookup; Pwned Passwords API (k-anonymity).
- Dehashed — credential search (paid).
- IntelX — data intelligence.
- LeakCheck, Snusbase, BreachDirectory, Scattered Secrets, Phonebook, LeakPeek.
- Cavalier (Hudson Rock) — infostealer log lookups; FREE; highest single-source ROI for finding compromised employee credentials in corporate SSO.
15.0.1 HudsonRock Cavalier — direct API recipe
The web UI wraps a public, unauthenticated JSON API. Hit it directly:
# By domain (canonical first call)
curl -sk -m 30 "https://cavalier.hudsonrock.com/api/json/v2/osint-tools/search-by-domain?domain=target.com" | jq .
# By email (single-account check)
curl -sk -m 30 "https://cavalier.hudsonrock.com/api/json/v2/osint-tools/search-by-email?email=alice@target.com" | jq .
# By URL (when target's app is the breach victim)
curl -sk -m 30 "https://cavalier.hudsonrock.com/api/json/v2/osint-tools/search-by-url?url=https://app.target.com" | jq .
PowerShell:
$hr = Invoke-RestMethod -Uri "https://cavalier.hudsonrock.com/api/json/v2/osint-tools/search-by-domain?domain=$D" -TimeoutSec 30
"Employees: $($hr.employees) | Users: $($hr.users) | Third-party: $($hr.third_parties) | Total: $($hr.total)"
$hr.data.employees_urls | Sort-Object -Property occurrence -Descending | Select-Object -First 20
$hr.data.clients_urls | Sort-Object -Property occurrence -Descending | Select-Object -First 15
Top-level JSON fields:
total— total stealer entries touching this domain.totalStealers— global stealer-log corpus size (context only).employees— count of<*>@<domain>accounts found.users— count of accounts where the domain appeared as a visited URL (customers/vendors).third_parties— accounts touching adjacent domains in the org.data.employees_urls[]—{occurrence, type, url}— internal apps where employees were logging in when stolen. Subdomain hits here = recon gold.data.clients_urls[]— same shape; user-facing apps (often reveals undocumented public portals).data.stealer_families[]—{_key, _value}→ which stealer (RedLine / Lumma / StealC / Vidar / Raccoon).data.dates_compromised[]—{_key, _value}→ temporal distribution.
Free-tier caveats (CRITICAL to know):
- Subdomain hostnames in
data.*_urls[]past the first few are redacted with asterisks (*****.target.com). Pivot to paid Cavalier tier or other sources for unredacted. - Free endpoint returns counts + sample URLs only. Cleartext passwords + emails are never in the free response.
- Rate limit ~1 req/sec/IP; 429 on burst. Sleep 1s between calls.
- For unredacted creds + bulk enumeration → paid Cavalier portal.
Severity mapping (per §15.1 + §15.2): employees ≥ 10 → CRITICAL, regardless of whether the breached service is still online (legacy Lotus Domino / on-prem mail decommissioned + cloud SSO migration → employees almost always reuse passwords → SSO_EXPOSURE escalates CRITICAL).
15.1 Domain-Level Breach Severity Mapping
When you query a breach corpus by domain, map the result to severity like so:
| Stat | Severity |
|---|---|
| ≥ 10 employees compromised | CRITICAL |
| 1–9 employees compromised | HIGH |
| ≥ 1 end-user (non-employee) compromised | MEDIUM |
| Domain seen in breach with 0 named accounts | INFO |
Employees vs end-users distinction: an employee account is <anything>@<target-domain> (the breach victim is the target's own staff). An end-user account is the target's customer who reused a password — useful for credential-stuffing risk awareness but not directly compromising the target's identity fabric.
15.2 SSO_EXPOSURE finding
When a discovered SSO tenant (Entra GUID / Okta slug / Google Workspace domain) intersects with the breach corpus on its domain → SSO_EXPOSURE finding, severity CRITICAL. Evidence: tenant ID + product + employee count + per-account source attribution.
Legacy-mail-decommissioned pattern (high-value variant):
If mail.<domain> / webmail.<domain> returns NXDOMAIN today but HudsonRock/HIBP corpus still has historical employee credentials against it AND autodiscover.<domain> resolves to Microsoft IPs (M365) or aspmx.l.google.com MX (Workspace), the org migrated from on-prem to cloud — and the stolen passwords almost certainly survived the migration via password reuse. Escalate to CRITICAL SSO_EXPOSURE even when the legacy host is dead.
Concrete triggers (all three together):
Resolve-DnsName mail.<domain> -Type A→ NXDOMAIN (legacy gone)- HudsonRock corpus has employee URLs against the old host (e.g.
mail.<domain>/names.nsffor Lotus Domino,mail.<domain>/owa/for Exchange,mail.<domain>/iwaredir.nsffor iNotes,mail.<domain>/zimbra/for Zimbra) - Current MX → M365 / Google Workspace / Zoho cloud (DNS confirms migration)
Evidence pack: tenant GUID + breach count + 3+ legacy URLs from corpus + autodiscover Microsoft IPs + current MX. Recommend forced password rotation + MFA audit + Conditional Access review.
16. Pre-built Wordlists & Probe Paths
Copy-pasteable arsenals, severity-annotated where relevant.
16.1 Swagger / OpenAPI discovery — 28 paths
Probe each path on every alive webapp. GET (or HEAD if rate-limited).
swagger.json
swagger.yaml
swagger/v1/swagger.json
swagger/v2/swagger.json
swagger-ui.html
swagger-ui/
swagger-resources
api-docs
api-docs.json
api/swagger
api/swagger.json
api/swagger-ui.html
api/v1/swagger.json
api/v2/swagger.json
api/v3/api-docs
v2/api-docs
v3/api-docs
openapi.json
openapi.yaml
openapi/v1
openapi/v3
docs
redoc
rapidoc
api/docs
api/documentation
.well-known/openapi
Severity:
- Reachable Swagger/OpenAPI spec without auth → HIGH
LEAKY_API_SPEC(full endpoint enumeration leaks; often reveals undocumented internal APIs). - Behind auth but accessible to any authenticated user → MEDIUM (still discloses internal API surface).
16.2 GraphQL discovery — 13 paths
graphql
graphiql
api/graphql
v1/graphql
v2/graphql
query
api/query
gql
altair
playground
subscriptions
graphql/console
api/v1/graphql
Standard introspection POST body:
{
"operationName": "IntrospectionQuery",
"query": "query IntrospectionQuery { __schema { types { name kind fields { name type { name kind } } } queryType { name } mutationType { name } subscriptionType { name } } }"
}
Severity:
- Introspection returns schema without auth → HIGH
OPEN_GRAPHQL_API. - Field-suggestion enumeration possible (server returns "did you mean" for typo'd field names) → MEDIUM (re-derive partial schema even when introspection is disabled).
/graphqlaccepts batched queries ([...]request body) → MEDIUM (rate-limit bypass surface; auth bypass via mixed batches).
UI markers (lower severity but still discoverable):
- HTML response contains
graphiql,playground,apollo studio,altair→ GraphiQL UI exposed (often shipped accidentally on prod).
16.3 High-risk ports — 35 services
For each open port, emit a finding with the severity and "why an attacker cares" below. Source for the open-port observation: Shodan InternetDB (free, 1 req/sec) is the recommended starting point.
| Port | Service | Severity | Why it matters |
|---|---|---|---|
| 21 | FTP | HIGH | Anonymous read often enabled; cleartext creds. |
| 22 | SSH | LOW | Banner discloses version; brute-force surface. |
| 23 | Telnet | HIGH | Cleartext protocol; should never be exposed. |
| 25 | SMTP | LOW | Open relay risk; version banner. |
| 53 | DNS | LOW | Recursion = DDoS amplifier; AXFR opportunism. |
| 80 | HTTP | INFO | Standard. |
| 110 | POP3 | LOW | Cleartext if no STARTTLS. |
| 111 | rpcbind | MEDIUM | NFS exports enumeration. |
| 135 | MS RPC | HIGH | Enum via Impacket. |
| 139 | NetBIOS-SSN | HIGH | File/printer enum. |
| 143 | IMAP | LOW | Cleartext if no STARTTLS. |
| 161 | SNMP | HIGH | Community strings often public/private; full device enum. |
| 389 | LDAP | HIGH | Anonymous bind = full directory dump. |
| 443 | HTTPS | INFO | Standard. |
| 445 | SMB | CRITICAL | EternalBlue, SMB relay, anonymous shares. |
| 465 | SMTPS | LOW | Banner. |
| 514 | rsyslog | MEDIUM | Log injection / DoS. |
| 587 | SMTP-MSA | LOW | Banner. |
| 631 | IPP/CUPS | MEDIUM | Print server enum / RCE in old CUPS. |
| 873 | rsync | HIGH | Modules often listable; backup data exposure. |
| 1433 | MSSQL | HIGH | Brute-force; xp_cmdshell. |
| 1521 | Oracle TNS | HIGH | Brute-force; SID enum. |
| 2049 | NFS | HIGH | World-readable exports. |
| 2375 | Docker API (unencrypted) | CRITICAL | Unauthenticated container/host takeover. |
| 2376 | Docker API (TLS) | HIGH | Cert validation bypass risk. |
| 3000 | Common dev / Grafana | MEDIUM | Often Grafana / Express dev with default creds. |
| 3306 | MySQL | HIGH | Brute-force; default root:"". |
| 3389 | RDP | CRITICAL | BlueKeep / DejaBlue / NLA bypass. |
| 5432 | PostgreSQL | HIGH | Brute-force; default postgres:postgres. |
| 5601 | Kibana | HIGH | Often unauthenticated; Elasticsearch pivot. |
| 5900 | VNC | HIGH | Often unauthenticated or weak password. |
| 5984 | CouchDB | HIGH | Default no auth; admin party. |
| 6379 | Redis | CRITICAL | No auth default; write authorized_keys for SSH. |
| 7001 | WebLogic | HIGH | Frequent CVEs (CVE-2020-14882, etc.). |
| 8000 | Common dev | MEDIUM | Django, common dev servers. |
| 8080 | HTTP-alt | MEDIUM | Tomcat, Jenkins, common proxy. |
| 8443 | HTTPS-alt | MEDIUM | Same as 8080. |
| 8888 | Common dev / Jupyter | HIGH | Jupyter often exposes interactive shell. |
| 9090 | Cockpit / Prometheus | HIGH | Server admin UI / metrics scraping. |
| 9200 | Elasticsearch | CRITICAL | Typically no auth. |
| 9300 | Elasticsearch transport | HIGH | Cluster join + RCE. |
| 11211 | memcached | MEDIUM | UDP DDoS amp; data dump. |
| 27017 | MongoDB | CRITICAL | No auth by default. |
| 50070 | Hadoop NameNode | HIGH | HDFS browse. |
When Shodan InternetDB returns vulns[] for a port, escalate the finding severity by one tier and include the CVE list in evidence.
16.4 Missing security headers — 6 findings
For every alive webapp, audit response headers. Each missing header below = one finding.
| Header | Severity (default) | Severity (sensitive path) | Notes |
|---|---|---|---|
Strict-Transport-Security |
MEDIUM | HIGH | Sensitive paths: /login, /signin, /sso, /admin, /auth. |
Content-Security-Policy |
MEDIUM | MEDIUM | XSS impact mitigation gone. |
X-Frame-Options |
LOW | LOW | Clickjacking. (CSP frame-ancestors is the modern replacement.) |
X-Content-Type-Options |
LOW | LOW | MIME-sniff XSS. |
Referrer-Policy |
INFO | INFO | Outbound link leakage. |
Permissions-Policy |
INFO | INFO | Feature-policy hardening. |
16.5 Always-on HTTP checks — 15 paths
Run these against every alive webapp regardless of Nuclei availability. Cheap; high signal.
| Path | Finding | Severity | Match logic |
|---|---|---|---|
/.git/config |
Exposed .git repo |
CRITICAL | Body contains [core], [remote, repositoryformatversion |
/.git/HEAD |
Exposed .git/HEAD |
HIGH | Body matches ^ref:\s |
/.env |
Exposed .env |
CRITICAL | Multiline regex ^\s*[A-Z_][A-Z0-9_]*\s*= |
/server-status |
Apache server-status | MEDIUM | Body contains Apache Server Status or matching title |
/server-info |
Apache mod_info | MEDIUM | Body contains Apache Server Information |
/.DS_Store |
Exposed .DS_Store |
LOW | Byte signature \x00\x00\x00\x01Bud1 |
/phpinfo.php |
phpinfo() leak | HIGH | Body contains phpinfo(), PHP Version, or matching title |
/info.php |
phpinfo() (alt path) | HIGH | Same as above |
/actuator/env |
Spring Boot /actuator/env |
CRITICAL | Body contains "propertySources", systemProperties, systemEnvironment |
/actuator/heapdump |
Spring Boot heapdump | CRITICAL | HPROF magic bytes / large binary download |
/_cat/indices |
Elasticsearch open | HIGH | Returns index list |
/console |
Jenkins script console | HIGH | Body contains Jenkins/Script Console |
/manager/html |
Tomcat Manager | HIGH | Body contains Tomcat Web Application Manager |
/wp-admin/install.php |
Orphaned WP install | LOW | Body contains WordPress Installation |
/.well-known/security.txt |
Disclosure policy info | INFO | Parse contact + policy fields |
Plus parse /robots.txt for Disallow: paths — those become the next-tier wordlist for that target.
16.6 SAML metadata — 5 paths
/saml/metadata
/FederationMetadata/2007-06/FederationMetadata.xml
/federationmetadata/2007-06/federationmetadata.xml
/simplesaml/saml2/idp/metadata.php
/auth/saml2/metadata
Reachable SAML metadata XML reveals: EntityID, signing certs (often pinned → cert-reuse pivot), SingleSignOnService URL, NameIDFormat. Mark as MISCONFIG (LOW severity unless metadata leaks internal hostnames or non-public certs, then MEDIUM).
16.7 SSO subdomain prefixes — 8 prefixes
Probe each against root domain + every sibling brand domain:
auth.{domain}
login.{domain}
sso.{domain}
idp.{domain}
iam.{domain}
identity.{domain}
accounts.{domain}
oauth.{domain}
Plus probe /.well-known/openid-configuration on every alive subdomain (regardless of prefix).
16.8 Cloud bucket permutation arsenal
6 prefixes:
"" # bare candidate
backup-
assets-
static-
dev-
prod-
15 suffixes:
"" # bare candidate
-backup
-assets
-static
-media
-data
-uploads
-dev
-prod
-staging
-logs
-private
-public
-dump
-archive
47 generic stems (filter unless combined with target-identifying token):
www, mail, email, app, apps, web, webmail, ftp, cdn, static, assets, media, img, images,
videos, download, downloads, upload, uploads, data, files, docs, support, help, kb,
blog, news, dev, test, staging, stg, qa, uat, sandbox, preprod, preview, vpn,
mx, smtp, imap, pop, dns, ns, ns1, ns2, mx1, mx2
Provider URL templates:
S3:
https://{candidate}.s3.amazonaws.com/
https://{candidate}.s3-{region}.amazonaws.com/ # try us-east-1, us-west-2, eu-west-1, ap-southeast-1 first
https://s3.{region}.amazonaws.com/{candidate}/
GCS:
https://{candidate}.storage.googleapis.com/
https://storage.googleapis.com/{candidate}/
Azure Blob:
https://{candidate}.blob.core.windows.net/
Probe technique: HEAD first → 200/301 = exists, 403 = exists private, 404 = skip. On exists, GET root → if XML/JSON object listing returns, CRITICAL PUBLIC_CLOUD_BUCKET. Direct-URL object reads but not listable → HIGH PUBLIC_CLOUD_BUCKET_OBJECT_READ.
16.9 JS guess-paths for endpoint discovery
Probe these paths on every alive webapp (in addition to scraped <script src=...>):
/main.js
/app.js
/bundle.js
/runtime.js
/index.js
/vendor.js
/_next/static/_buildManifest.js
/_next/static/_ssgManifest.js
/static/js/main.js
/static/js/bundle.js
/assets/index.js
/static/js/main.<hash>.js # try hash discovery via 404 patterns
For every found JS, also try <jsfile>.map for sourcemap leaks (HIGH INFO_DISCLOSURE).
16.10 Endpoint extraction regex tiers
Three tiers, run in order on every JS body + every sourcesContent[] blob:
Tier 1 — generic quoted paths:
['"`](/[A-Za-z0-9_\-./{}\[\]?=&%:]+)['"`]
Match group: the path. High recall, lots of false positives — apply allowlist downstream.
Tier 2 — API-ish paths (biased filter on tier 1):
['"`](/(?:api|graphql|gql|v\d+|swagger|openapi|rest|services|internal|admin|auth|oauth|user|users|account|accounts|search|export|upload|file|files|download|webhook|hooks|callback|admin)/[A-Za-z0-9_\-./{}\[\]?=&%:]+)['"`]
Tier 3 — fully-qualified URLs:
\bhttps?://[A-Za-z0-9.\-]+\.[A-Za-z]{2,}(?::\d+)?[/A-Za-z0-9_\-./{}\[\]?=&%:#]*
Dedup on (method, normalized-path-template) where the template replaces /123/ with /{id}/ etc.
16.11 Internal-host leakage regexes
Run on every JS body + sourcesContent + APK strings + manifest:
RFC1918:
\b(?:10\.(?:\d{1,3}\.){2}\d{1,3}|172\.(?:1[6-9]|2\d|3[01])\.(?:\d{1,3})\.(?:\d{1,3})|192\.168\.(?:\d{1,3})\.(?:\d{1,3})|127\.(?:\d{1,3}\.){2}\d{1,3})\b
Internal DNS suffixes:
\b[A-Za-z0-9][A-Za-z0-9\-]{0,62}\.(?:internal|corp|lan|intranet|local|prod|staging|dev|qa|test)\b
Kubernetes service DNS:
\b[A-Za-z0-9\-]+\.[A-Za-z0-9\-]+\.svc(?:\.cluster\.local)?\b
Each match → MEDIUM INFO_DISCLOSURE. Aggregate per host: if many matches share the same internal subdomain, that's a recon seed for any future internal phase.
16.12 Subdomain-takeover provider fingerprints (summary, 27 providers)
Watch for these CNAME targets + the corresponding "available for claim" response signature:
| Provider | CNAME pattern | Takeover signature |
|---|---|---|
| GitHub Pages | *.github.io |
There isn't a GitHub Pages site here. |
| Heroku | *.herokuapp.com |
No such app |
| AWS S3 | *.s3*.amazonaws.com |
NoSuchBucket |
| AWS CloudFront | *.cloudfront.net |
Bad request w/ specific X-Amz error |
| Azure (multiple) | *.azurewebsites.net, *.blob.core.windows.net, *.cloudapp.net, *.trafficmanager.net |
Various per-product 404 patterns |
| Shopify | shops.myshopify.com |
Sorry, this shop is currently unavailable. |
| Squarespace | *.squarespace.com |
No Such Account |
| Tumblr | *.tumblr.com |
Whatever you were looking for doesn't currently exist. |
| WordPress | *.wordpress.com |
Do you want to register *.wordpress.com? |
| Fastly | various | Fastly-specific 404 |
| Pantheon | *.pantheonsite.io |
The gods are wise, but do not know of the site... |
| Surge.sh | *.surge.sh |
project not found |
| Bitbucket Pages | *.bitbucket.io |
Repository not found |
| Tilda | *.tilda.ws |
Please renew your subscription |
| Strikingly | *.s.strikinglydns.com |
PAGE NOT FOUND |
| Smartling | *.smartling.com |
Domain is not configured |
| Ngrok | *.ngrok.io |
Tunnel not found |
| Webflow | *.webflow.io |
Site not found |
| Zendesk | *.zendesk.com |
Help Center Closed |
| Cargo | *.cargocollective.com |
404 Not Found (with cargo branding) |
| Statuspage | *.statuspage.io |
Not found |
| Intercom | *.intercom.help |
Not found |
| Helpjuice | *.helpjuice.com |
Not found |
| Helpscout | *.helpscoutdocs.com |
Not found |
| Tictail | *.tictail.com |
Not found |
| Brightcove | *.brightcovegallery.com |
Not found |
| Smugmug | various | Not found |
For full per-provider detection signatures + edge cases, use SubdomainX or Subzy/Subjack against a freshly-fetched fingerprint database.
16.13 Copy-Paste Probes (curl one-liners)
Every probe path in §16.1–16.12 with a runnable curl. Defaults: -sk (silent + ignore TLS errors), -m 10 (10s max), -o /tmp/r (response body to disk), -w '%{http_code}\n' (print status code), -A "Mozilla/5.0" (UA — change per persona).
Always-on HTTP checks (§16.5):
T="https://target.example"
# .git/config (CRITICAL)
curl -sk -m 10 "$T/.git/config" | grep -E '\[core\]|\[remote|repositoryformatversion'
# .git/HEAD (HIGH)
curl -sk -m 10 "$T/.git/HEAD" | grep -E '^ref:'
# .env (CRITICAL)
curl -sk -m 10 "$T/.env" | grep -E '^[[:space:]]*[A-Z_][A-Z0-9_]*[[:space:]]*='
# Apache /server-status (MEDIUM)
curl -sk -m 10 "$T/server-status" | grep -i 'Apache Server Status'
# Apache /server-info (MEDIUM)
curl -sk -m 10 "$T/server-info" | grep -i 'Apache Server Information'
# .DS_Store (LOW)
curl -sk -m 10 "$T/.DS_Store" -o /tmp/dsstore && file /tmp/dsstore | grep -i 'data'
# phpinfo.php (HIGH)
curl -sk -m 10 "$T/phpinfo.php" | grep -E 'phpinfo\(\)|PHP Version'
# info.php (HIGH)
curl -sk -m 10 "$T/info.php" | grep -E 'phpinfo\(\)|PHP Version'
# Spring Boot /actuator/env (CRITICAL)
curl -sk -m 10 "$T/actuator/env" | grep -E '"propertySources"|systemProperties|systemEnvironment'
# Spring Boot /actuator/heapdump (CRITICAL — saves binary; check size)
curl -sk -m 30 "$T/actuator/heapdump" -o /tmp/heap && file /tmp/heap | grep -i 'HPROF\|data'
# Elasticsearch open (HIGH)
curl -sk -m 10 "$T/_cat/indices?v"
# Jenkins script console (HIGH)
curl -sk -m 10 "$T/script" | grep -iE 'Jenkins|Script Console'
# Tomcat manager (HIGH)
curl -sk -m 10 "$T/manager/html" -w '%{http_code}\n' | tail -1 # 401 = present + auth-gated; 200 = no auth
# WordPress orphan installer (LOW)
curl -sk -m 10 "$T/wp-admin/install.php" | grep -i 'WordPress Installation'
# security.txt (INFO)
curl -sk -m 10 "$T/.well-known/security.txt"
SSO subdomain prefixes (§16.7):
D="target.example"
for prefix in auth login sso idp iam identity accounts oauth; do
echo "=== ${prefix}.${D} ==="
curl -sk -m 10 "https://${prefix}.${D}/.well-known/openid-configuration" -o /dev/null -w '%{http_code}\n'
done
# Generic OIDC discovery on any host:
curl -sk -m 10 "https://${HOST}/.well-known/openid-configuration" | jq .
SAML metadata paths (§16.6):
H="target.example.com"
for p in /saml/metadata \
/FederationMetadata/2007-06/FederationMetadata.xml \
/federationmetadata/2007-06/federationmetadata.xml \
/simplesaml/saml2/idp/metadata.php \
/auth/saml2/metadata; do
echo "=== $p ==="
curl -sk -m 10 "https://${H}${p}" -o /dev/null -w '%{http_code} %{size_download}\n'
done
Cloud bucket probes (§16.8):
B="candidate-bucket-name"
# S3 (us-east-1 first)
curl -sk -m 10 -I "https://${B}.s3.amazonaws.com/" -w 'STATUS:%{http_code}\n' | head -20
# If 200/301: list objects
curl -sk -m 10 "https://${B}.s3.amazonaws.com/?list-type=2" | head -50
# S3 region-specific
for r in us-east-1 us-west-2 eu-west-1 ap-southeast-1; do
curl -sk -m 10 -I "https://${B}.s3-${r}.amazonaws.com/" -w "${r}: %{http_code}\n"
done
# GCS
curl -sk -m 10 -I "https://${B}.storage.googleapis.com/"
curl -sk -m 10 "https://storage.googleapis.com/${B}/"
# Azure Blob
curl -sk -m 10 -I "https://${B}.blob.core.windows.net/"
curl -sk -m 10 "https://${B}.blob.core.windows.net/?comp=list"
GraphQL introspection POST (§16.2):
H="https://target.example/graphql"
curl -sk -m 15 -X POST "$H" \
-H 'Content-Type: application/json' \
-d '{
"operationName":"IntrospectionQuery",
"query":"query IntrospectionQuery { __schema { types { name kind fields { name type { name kind } } } queryType { name } mutationType { name } subscriptionType { name } } }"
}' | jq '.data.__schema.types | length'
Read-only secret validators (§23):
# Postman PMAK
curl -sk -m 10 -H "X-Api-Key: PMAK-..." https://api.getpostman.com/me | jq .
# AWS (use boto3 instead of curl — pre-signing complexity)
python3 -c "import boto3; print(boto3.client('sts', aws_access_key_id='AKIA...', aws_secret_access_key='...').get_caller_identity())"
# GitHub PAT (note scope header)
curl -sk -m 10 -H "Authorization: token ghp_..." https://api.github.com/user -D /tmp/h | jq -r '.login,.email'
grep -i 'X-OAuth-Scopes' /tmp/h
# Slack
curl -sk -m 10 -H "Authorization: Bearer xoxb-..." -X POST https://slack.com/api/auth.test | jq .
# Anthropic (read-only validation)
curl -sk -m 10 -H "x-api-key: sk-ant-..." -H "anthropic-version: 2023-06-01" https://api.anthropic.com/v1/models | jq '.data | length'
# OpenAI
curl -sk -m 10 -H "Authorization: Bearer sk-..." https://api.openai.com/v1/models | jq '.data | length'
# npm
curl -sk -m 10 -H "Authorization: Bearer npm_..." https://registry.npmjs.org/-/whoami | jq .
# Atlassian (account)
curl -sk -m 10 -u "email:ATATT3xFfGF0_..." https://your-domain.atlassian.net/rest/api/3/myself | jq .
# DataDog (API + APP key both required)
curl -sk -m 10 -H "DD-API-KEY: ..." -H "DD-APPLICATION-KEY: ..." https://api.datadoghq.com/api/v1/validate | jq .
Bulk webapp triage (httpx, faster than curl loop):
# Install: go install github.com/projectdiscovery/httpx/cmd/httpx@latest
echo "target.example" | httpx -sc -title -tech-detect -web-server -ip -cdn -follow-redirects
# With probe list
cat subdomains.txt | httpx -sc -title -tech-detect -path /actuator/env,/.git/config,/.env -mc 200,301,403
Save responses for evidence:
mkdir -p evidence/$(date -u +%Y%m%d)
T="https://target.example"
P="/actuator/env"
TS=$(date -u +%Y%m%dT%H%M%SZ)
SAFE_NAME=$(echo "${T}${P}" | tr '/:' '_')
curl -sk -m 10 "$T$P" -o "evidence/$(date -u +%Y%m%d)/${TS}_${SAFE_NAME}.body" \
-D "evidence/$(date -u +%Y%m%d)/${TS}_${SAFE_NAME}.headers"
sha256sum "evidence/$(date -u +%Y%m%d)/${TS}_${SAFE_NAME}".* > "evidence/$(date -u +%Y%m%d)/${TS}_${SAFE_NAME}.sha256"
16.14 Email Security Analysis (SPF/DMARC/DKIM/BIMI/MTA-STS/DNSSEC)
Spoof feasibility + SaaS tenant inference from a target's email DNS.
SPF lookup + parsing:
D="target.example"
dig +short TXT "$D" | grep -i 'v=spf1'
Common SPF parsing checklist:
- Ends in
-all(hardfail) → strict; major providers reject spoofs. - Ends in
~all(softfail) → spam folder for spoofs. - Ends in
?allor noall→ permissive; spoofs likely deliver. - Includes (
include:) reveal SaaS tenants:include:_spf.google.com→ Google Workspace.include:spf.protection.outlook.com→ Microsoft 365.include:_spf.salesforce.com→ Salesforce.include:mail.zendesk.com→ Zendesk customer.include:sendgrid.net→ SendGrid customer.include:mailgun.org→ Mailgun customer.include:_spf.atlassian.net→ Atlassian Cloud.include:amazonses.com→ AWS SES.include:mktomail.com→ Marketo.include:_spf.intuit.com→ Intuit (QuickBooks/Mailchimp).include:spf.mandrillapp.com→ Mandrill.include:_spf.workday.com→ Workday.
If SPF includes ≥10 mechanisms (max-lookups limit) → SPF eval likely fails → spoofs may pass. Tools: spfquery, spftools (online), dig +trace.
DMARC policy + alignment:
dig +short TXT "_dmarc.${D}"
Parse for:
p=→ primary policy (none,quarantine,reject).sp=→ subdomain policy (defaults top=).aspf=/adkim=→ alignment mode (r=relaxed,s=strict).pct=→ percentage of mail to which policy applies.rua=/ruf=→ reporting addresses (often reveals SaaS DMARC vendors: dmarcian, valimail, Agari, easydmarc).
Severity:
p=none→ spoof-feasible, downgrade trust → MEDIUM finding.p=quarantine pct<100→ partial enforcement → LOW.p=reject+aspf=s+adkim=s→ well-postured → no finding.
DKIM key discovery:
DKIM selectors aren't well-known; common patterns:
for selector in default google selector1 selector2 mail email k1 dkim s1 s2 mta1 mta2 \
amazonses 20240101 20230101 mailchimp sendgrid mxvault; do
echo "=== ${selector} ==="
dig +short TXT "${selector}._domainkey.${D}"
done
If a key returns: extract p=<base64> and check key length. RSA-1024 → MEDIUM (deprecated; should be 2048+). Missing or rotated infrequently → LOW finding.
BIMI (Brand Indicators for Message Identification):
dig +short TXT "default._bimi.${D}"
If present + p=reject DMARC → brand-impersonation defense in inbox UI. Absence is LOW only (operational, not exploitable).
MTA-STS (Mail Transfer Agent Strict Transport Security):
dig +short TXT "_mta-sts.${D}"
curl -sk -m 10 "https://mta-sts.${D}/.well-known/mta-sts.txt"
If neither responds → MX-server TLS not enforced; MITM-able. LOW finding. If mode=enforce present and policy file matches → well-postured.
TLS-RPT (TLS Reporting):
dig +short TXT "_smtp._tls.${D}"
DNSSEC validation:
dig +dnssec "${D}" SOA | grep -E 'flags|RRSIG'
delv "${D}" 2>&1 | grep -i 'fully validated\|insecur'
If delv returns "insecure" → DNSSEC not enabled (LOW finding; doesn't enable spoof but is hardening gap).
MX → IdP / mail-host inference:
dig +short MX "${D}"
| MX pattern | IdP / hosting |
|---|---|
aspmx.l.google.com, *.googlemail.com |
Google Workspace |
*.mail.protection.outlook.com |
Microsoft 365 |
*.mail.eo.outlook.com |
Microsoft 365 (older) |
*.zoho.com |
Zoho Mail |
*.yandex.net |
Yandex 360 |
*.fastmail.com |
Fastmail |
*.proofpoint.com, *.pphosted.com |
Proofpoint (M365 user with Proofpoint inbound) |
*.mimecast.com, *.mimecast-eu.com |
Mimecast |
*.barracudanetworks.com |
Barracuda |
| Self-hosted IPs in target ASN | On-prem mail server (often Exchange) |
DMARC reporting-vendor inference (parse rua= / ruf=):
| RUA/RUF host | Vendor | Implication |
|---|---|---|
*.dmarcian.com |
dmarcian | DMARC reporting customer |
*.valimail.com, *.dmarc-rua.com |
Valimail | DMARC reporting customer |
*.kdmarc.com |
Kratikal kDMARC | Indian DMARC vendor; common in IN orgs |
*.agari.com |
Agari (Fortra) | Email security vendor |
*.easydmarc.com |
EasyDMARC | DMARC reporting customer |
*.dmarcanalyzer.com |
DMARC Analyzer | Reporting customer |
*.postmarkapp.com |
Postmark | DMARC reporting addon |
| `@<targ |
…(truncated)