Analyzing Pdf Malware With Pdfid
Overview
Cybersecurity skill for analyzing pdf malware with pdfid. Follows industry best practices and security standards.
When to Use
Trigger phrases:
"analyzing pdf malware with pdfid"
"Analyzes malicious PDF files using PDFiD, pdf-parser, and peepdf to identify emb"
A suspicious PDF attachment has been flagged by email security or reported by a user
You need to determine if a PDF contains embedded JavaScript, shellcode, or exploit code
Triaging PDF documents before opening them in a sandbox or analysis environment
Extracting embedded executables, scripts, or URLs from malicious PDF objects
Analyzing PDF exploit kits targeting Adobe Reader or other PDF viewer vulnerabilities
Do not use for analyzing the rendered visual content of a PDF; this is for structural analysis of the PDF file format for malicious objects.
When NOT to Use
- When you lack proper authorization for testing
- For production systems without change management
- When the task requires legal or compliance expertise beyond technical scope
Prerequisites
- Python 3.8+ with Didier Stevens' PDF tools installed (
pip install pdfid pdf-parser)
- peepdf installed for interactive PDF analysis (
pip install peepdf)
- pdftotext from poppler-utils for extracting text content safely
- YARA with PDF-specific rules for malware family identification
- Isolated analysis VM without a PDF reader installed (prevent accidental opening)
- CyberChef for decoding embedded Base64, hex, or deflate streams
Workflow
# Example: IOC detection
import re
IOC_PATTERNS = {
"ip": r"\b(?:\d{1,3}\.){3}\d{1,3}\b",
"domain": r"\b[a-z0-9-]+\.[a-z]{2,}\b",
"hash_md5": r"\b[a-f0-9]{32}\b",
"hash_sha256": r"\b[a-f0-9]{64}\b",
}
def extract_iocs(text: str) -> dict:
return {k: re.findall(v, text) for k, v in IOC_PATTERNS.items()}
- Scope the Analysis — Define what pdf malware artifacts or data sources to examine and the investigation timeline.
- Preserve Evidence — Create forensic copies of relevant data. Maintain chain of custody documentation.
- Extract Key Indicators — Use pdfid to parse and extract relevant pdf malware data points from collected artifacts.
- Correlate Findings — Cross-reference extracted data with other sources (threat intel, logs, timelines).
- Build Timeline — Construct a chronological sequence of events related to pdf malware.
- Document Analysis — Write findings report with evidence, conclusions, and recommendations.
Tools
- pdfid — Primary tool for this skill
- Forensic Toolkit — Evidence collection and analysis
- Timeline Tools — Chronological event reconstruction
- Log Analysis Platform — Centralized log parsing and search
Process
- Reconnaissance — Gather target information, identify attack surface, enumerate services
- Analysis/Exploitation — Execute the technique, analyze results, document findings
- Reporting — Document IOCs, write findings, provide remediation recommendations
Verification
Anti-Rationalization Table
| Rationalization |
Reality |
| "We are too small to be targeted" |
Automated attacks target everyone. Size does not matter. |
| "Security slows us down" |
A breach slows you down 100x more. Build security in from the start. |
| "We will fix it after launch" |
Vulnerabilities in production are exploited within hours. Fix before deploy. |
1---2name: analyzing-pdf-malware-with-pdfid3description: Use when analyzes malicious PDF files using PDFiD, pdf-parser, and peepdf to identify embedded JavaScript, shellcode, exploits, and suspicious objects without opening the document. Determines the attack vector and extracts embedded payloads for further analysis. Activates for requests involving PDF malware analysis, malicious document analysis, PDF exploit investigation, or suspicious attachment triage. . Use when working with analyzing pdf malware with pdfid.4license: Apache-2.05---67# Analyzing Pdf Malware With Pdfid89## Overview1011Cybersecurity skill for analyzing pdf malware with pdfid. Follows industry best practices and security standards.1213## When to Use14**Trigger phrases:**15- "analyzing pdf malware with pdfid"16- "Analyzes malicious PDF files using PDFiD, pdf-parser, and peepdf to identify emb"171819- A suspicious PDF attachment has been flagged by email security or reported by a user20- You need to determine if a PDF contains embedded JavaScript, shellcode, or exploit code21- Triaging PDF documents before opening them in a sandbox or analysis environment22- Extracting embedded executables, scripts, or URLs from malicious PDF objects23- Analyzing PDF exploit kits targeting Adobe Reader or other PDF viewer vulnerabilities2425**Do not use** for analyzing the rendered visual content of a PDF; this is for structural analysis of the PDF file format for malicious objects.262728## When NOT to Use2930- When you lack proper authorization for testing31- For production systems without change management32- When the task requires legal or compliance expertise beyond technical scope333435## Prerequisites3637- Python 3.8+ with Didier Stevens' PDF tools installed (`pip install pdfid pdf-parser`)38- peepdf installed for interactive PDF analysis (`pip install peepdf`)39- pdftotext from poppler-utils for extracting text content safely40- YARA with PDF-specific rules for malware family identification41- Isolated analysis VM without a PDF reader installed (prevent accidental opening)42- CyberChef for decoding embedded Base64, hex, or deflate streams4344## Workflow4546```python47# Example: IOC detection48import re4950IOC_PATTERNS = {51 "ip": r"\b(?:\d{1,3}\.){3}\d{1,3}\b",52 "domain": r"\b[a-z0-9-]+\.[a-z]{2,}\b",53 "hash_md5": r"\b[a-f0-9]{32}\b",54 "hash_sha256": r"\b[a-f0-9]{64}\b",55}5657def extract_iocs(text: str) -> dict:58 return {k: re.findall(v, text) for k, v in IOC_PATTERNS.items()}59```60611. **Scope the Analysis** — Define what pdf malware artifacts or data sources to examine and the investigation timeline.622. **Preserve Evidence** — Create forensic copies of relevant data. Maintain chain of custody documentation.633. **Extract Key Indicators** — Use pdfid to parse and extract relevant pdf malware data points from collected artifacts.644. **Correlate Findings** — Cross-reference extracted data with other sources (threat intel, logs, timelines).655. **Build Timeline** — Construct a chronological sequence of events related to pdf malware.666. **Document Analysis** — Write findings report with evidence, conclusions, and recommendations.6768## Tools6970- **pdfid** — Primary tool for this skill71- **Forensic Toolkit** — Evidence collection and analysis72- **Timeline Tools** — Chronological event reconstruction73- **Log Analysis Platform** — Centralized log parsing and search747576## Process77781. **Reconnaissance** — Gather target information, identify attack surface, enumerate services791. **Analysis/Exploitation** — Execute the technique, analyze results, document findings801. **Reporting** — Document IOCs, write findings, provide remediation recommendations8182## Verification8384- [ ] All pdf malware procedures executed completely and documented85- [ ] Findings validated against multiple data sources86- [ ] False positives identified and filtered87- [ ] Results documented with evidence and timestamps88- [ ] Recommendations provided with risk-based prioritization8990## Anti-Rationalization Table9192| Rationalization | Reality |93|---|---|94| "We are too small to be targeted" | Automated attacks target everyone. Size does not matter. |95| "Security slows us down" | A breach slows you down 100x more. Build security in from the start. |96| "We will fix it after launch" | Vulnerabilities in production are exploited within hours. Fix before deploy. |