SKILL: Fuzzing
Metadata
Description
Practical fuzzing methodology reference: target identification, fuzzer selection (AFL++, libFuzzer, Boofuzz, Peach), harness writing, corpus curation, mutation strategies, coverage measurement, and crash triage. Use when setting up or running fuzz campaigns against any target.
Trigger Phrases
Use this skill when the conversation involves any of:
fuzzing, AFL++, libFuzzer, Boofuzz, Peach, fuzz harness, corpus, mutation, coverage, crash triage, network fuzzing, file format fuzzing
Instructions for Claude
When this skill is active:
- Load and apply the full methodology below as your operational checklist
- Follow steps in order unless the user specifies otherwise
- For each technique, consider applicability to the current target/context
- Track which checklist items have been completed
- Suggest next steps based on findings
Full Methodology
Fuzzing
automated software testing technique that is used for program analysis
Types
BlackBox
- Generates random or model‑based inputs and feeds them into the program under test
- Fuzzer has no knowledge of the program internals during fuzzing
- Pros
- Extremely fast
- Easy to use
- Scalable
- Cons
- Random testing leads to poor coverage
GreyBox
- Relies on lightweight instrumentation of the program under test
- Fuzzer has some knowledge of the program internals during fuzzing
- Pros
- Feedback
- Scalable
- Relatively fast
- Cons
- Examples
- General: AFL++, Honggfuzz, BooFuzz, WinAFL
- Directed: UAFuzz, AFLGo
- Grammar Based: Tlspuffin, AFLSmart
- Taint Based: Angora,
- Concolic: QSYM, AFL, SymSan
- Self‑learning: SieveFuzz, RL‑Fuzz
- Binary Fuzzing: Jackalope
- Special Fuzzers: Papora, Medusa
- Kernel: syzkaller, kAFL, wtf, KF/x
- ETC: LibAFL (snapshot & Unicorn support 0.15.x), FuzzBench, ClusterFuzz, DynamoRIO
Snapshot
- Takes and restores memory/register snapshots of the target or VM to bypass expensive initialization.
- Pros
- Order‑of‑magnitude faster execution on large or stateful binaries
- Deterministic and easily reproducible
- Works with closed‑source targets (no rebuild needed)
- Cons
- Initial snapshot & harness creation can be complex
- Requires specialised snapshot engine support (e.g., Nyx, Snapchange, wtf)
- Examples
- Nyx (v2 with Intel‑PT support; note: Nyx integration requires the AFL++-Nyx fork, not mainline AFL++ 4.x)
- Snapchange
- wtf
Snapshot Fuzzing Recipes
NYX_MODE=1 AFL_MAP_SIZE=1048576 afl-fuzz -i seeds -o findings -- ./target_nyx @@
WhiteBox
- Leverages heavyweight symbolic execution and static analysis, assuming full source availability
- Pros
- Solves path constraints to get past difficult logic
- Capable of generating inputs for deep, hard‑to‑reach program states
- Cons
- Complex
- Slow
- Not scalable
Ensemble Fuzzing
Ensemble fuzzing coordinates multiple heterogeneous fuzzers sharing a unified corpus, achieving better coverage than any single fuzzer.
Concept
- Cross-Pollination: Different fuzzers discover different code paths; sharing corpus maximizes coverage
- Complementary Strengths: AFL++ excels at control flow; Honggfuzz at crash detection; libFuzzer at in-process fuzzing
- Coordinated Scheduling: Use reinforcement learning or heuristics to select optimal fuzzer per input
Tools and Frameworks
- EnFuzz: Uses reinforcement learning to dynamically select which fuzzer runs on each input based on historical effectiveness
- CollAFL: Lightweight synchronization protocol for parallel fuzzer coordination without central coordinator
- CUPID: Analyzes feedback from all fuzzers to guide mutation strategies across the ensemble
Practical Setup
# Terminal 1: AFL++ primary with shared sync directory
afl-fuzz -M fuzzer1 -i seeds -o sync_dir -- ./target @@
# Terminal 2: Honggfuzz syncing to AFL corpus
../honggfuzz/honggfuzz -i sync_dir/fuzzer1/queue -W sync_dir/hfuzz \
--linux_perf_ipt_block -t 10 -- ./target ___FILE___
# Terminal 3: libFuzzer reading from both corpora
./target_libfuzzer sync_dir/fuzzer1/queue sync_dir/hfuzz/queue \
-max_total_time=3600 -runs=1000000
# Terminal 4: Sync monitor (optional)
watch -n 60 'ls -lh sync_dir/*/queue | wc -l'
Best Practices
- Start Simple: Begin with AFL++ + Honggfuzz; add libFuzzer if you have source
- Corpus Deduplication: Run
afl-cmin periodically to minimize combined corpus
- Resource Allocation: Allocate 40% AFL++, 30% Honggfuzz, 30% libFuzzer based on empirical results
- Monitoring: Track unique crashes per fuzzer to identify which contributes most
Performance Gains
- Research shows 15-40% more coverage vs single fuzzer
- 2-3× more unique crashes in 24-hour runs
- Particularly effective on complex parsers with multiple code paths
Components
Power Scheduler
- Responsible for distributing the amount of fuzzing time among the seeds in the seed queue
- Prioritize more interesting seeds
- Usually measure by its ability to produce new code coverage
- Assigns a scope to each seed and picks the one with highest score
- Self‑learning schedulers such as SieveFuzz and RL‑Fuzz use reinforcement learning to allocate energy dynamically based on past mutation success.
- Explore/Exploit auto‑tuning: AFL++ 4.30+ folds MOpt into the core and dynamically balances energy between seeds.
Mutation
- Responsible for making small changes to the fuzzing seed, with the goal of exercising new program behavior
- Examples: bit-flips, addition/subtraction, deletion, etc
- Typically operate by making small, incremental changes
- LLM‑assisted mutations & harness scaffolding: Large‑language models can infer grammars/dictionaries (
--dict2) and draft starter harness code (e.g., LibAFL's llm_mutator).
- Overflow‑aware arithmetic heuristics: Enable
--arith8/16/32-smart in AFL++ (or equivalents) to target boundary‑condition overflows.
- AFLPlusPlus
- Deterministic: includes single deterministic mutations on the content of the test cases
- Havoc: mutations are randomly stacked and also includes changes to the size of the test case
- Note: honggfuzz and libFuzzer implement weighted random mutation stacks instead of the deterministic/havoc split
- cmp_log guided mutations: Enable
-cmp_log (AFL++) or AFL_LLVM_CMPLOG=1 to perform Redqueen‑style input‑to‑state transforms.
- Snapshot‑aware mutations: With Nyx mode (
NYX_MODE=1) AFL++ randomises in‑memory state as well as file input for deep state machines.
- LLM‑guided edits: HyLLfuzz uses GPT/Llama‑3 to generate branch‑targeted mutations, achieving ≈ 1.3× edge coverage.
Directed Fuzzing (reach a specific bug/function)
- AFLGo basics:
- libAFL targeted stages: use objective feedbacks to push towards addresses/symbols of interest
Executor
- Responsible for executing the program under test with mutated fuzzing seed in a performant manner
- AFL++ uses a forkserver to skip the cost of expensive initialization and startup code
- AFL++ also offers persistent fuzzing mode, where the essential code is run inside a loop without a need for fork
- GPU‑accelerated execution loops are possible with CUDA‑AFL and LibAFL's
VulkanExecutor, ideal for shader or graphics targets.
- Recent AFL++ versions ship a Nyx executor capable of high exec/s on user‑mode targets; activate with
NYX_MODE=1 (verify version and build flags).
Persistent Mode Implementation
HF_ITER Style (Honggfuzz):
Honggfuzz offers HF_ITER-style persistent mode that's often easier to implement than ASAN-style for complex targets:
#include "honggfuzz.h"
int main(int argc, char** argv) {
// Expensive initialization (run once)
initialize_target();
// Persistent fuzzing loop
for (;;) {
size_t len;
uint8_t *buf;
// Get input from honggfuzz
HF_ITER(&buf, &len);
// Convert to target-appropriate format
FILE* input_stream = fmemopen(buf, len, "r");
// Execute target with fuzzed input
int result = target_function(input_stream);
// Cleanup for next iteration
fclose(input_stream);
reset_target_state();
}
}
Memory Management Considerations:
// Prevent memory leaks in persistent mode
void reset_target_state() {
// Free any allocated memory
cleanup_buffers();
// Reset global state
memset(&global_state, 0, sizeof(global_state));
// Close any open file handles
close_handles();
}
ASAN-Style Alternative:
extern "C" int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) {
// Direct fuzzing interface - no manual loop needed
return target_function(data, size);
}
Feedback
- Set of program state information from running with a particular input seed
- Fuzzers use feedback to determine fitness of a given seed(basic block coverage, edge coverage, path coverage)
- AFLPlusPlus uses edge coverage (keeps a shared bitmap between the target and the fuzzer)
- AFL++ uses edge coverage (keeps a shared bitmap between the target and the fuzzer)
- Hardware trace sources such as Intel® Processor Trace (IPT) and ARM ETM/CoreSight provide near‑zero‑overhead edge coverage for large binaries.
- LibAFL 0.15.2 can derive edge coverage from Last Branch Records, giving zero‑instrumentation tracing on recent Intel CPUs (
LBRFeedback).
Intel PT Coverage Setup
Honggfuzz with Intel PT:
Intel PT provides hardware-based coverage tracking ideal for binary-only targets where source instrumentation isn't possible:
# Basic Intel PT setup with honggfuzz
../honggfuzz/honggfuzz -i input_corpus/ -W workspace/ --linux_perf_ipt_block -t 10 -- ./target @@
# Performance tuning
../honggfuzz/honggfuzz -i input_corpus/ -W workspace/ \
--linux_perf_ipt_block \
--linux_perf_ignore_unknown \
--threads 4 \
-t 30 -- ./target @@
Intel PT Requirements:
- Modern Intel CPU with PT support (check with
cat /proc/cpuinfo | grep intel_pt)
- Linux kernel with PT enabled (
CONFIG_INTEL_PT=y)
- Sufficient privileges or
CAP_SYS_ADMIN capability
Troubleshooting Intel PT:
# Check PT availability
dmesg | grep -i "intel.*pt"
cat /sys/devices/intel_pt/type
# Permission issues
echo 'kernel.perf_event_paranoid = -1' >> /etc/sysctl.conf
sysctl -p
# Alternative: use capsh for specific capability
capsh --caps="cap_sys_admin+eip" -- -c "./honggfuzz_command"
Oracle
- Detects if any interesting behavior occurred during execution of the program under test with a given input
- Code Sanitizers are special compile-time instrumentation for the program under test that perform run-time checks
- AFL++ offers
QASAN to run binaries that are not instrumented with ASan under QEMU with the AFL++ instrumentation
- Different Bugs Require Different Oracles
- No Source: Statically rewrite the binary or
YOLO it
- Decoupled Sanitization: use SAND to off-load sanitizer checks and speed up fuzzing
- Memory Safety: use Address or Leak or Memory Sanitizer (try HWASAN for lower memory usage); or use ARM's MTE for hardware‑assisted tagging on AArch64
- Concurrency Issues: use Thread Sanitizer – LLVM 18 adds detection for C++20 std::atomic_wait
- Type Safety: use Type Sanitizer
- Undefined Behavior: use
UBSan
- Heap Hardening: Scudo Hardened Allocator (part of LLVM's compiler-rt, available in recent Clang versions) detects overflows and double‑frees with minimal performance cost.
- Kernel Undefined Behavior: Kernel UBSan (KUBSan) — enable with
CONFIG_UBSAN_TRAP=y (verify availability in your kernel).
- Control‑Flow Integrity: KCFI (Clang 18) adds low-overhead forward-edge CFI; enable with
-fsanitize=kcfi.
- Data‑Flow: Data‑Flow Sanitizer v2 adds taint‑tracking (
-fsanitize=dataflow) and now works with libFuzzer.
- Security‑property oracles: Idempotency or differential fuzzing for APIs (e.g., DiffSink for compiler targets).
- Kernel Stuff: use
KASAN, KMSAN, KCSAN
- Logic Bugs: Differential or Property(Correctness-Idempotency) Oracles
Property/Differential Oracles (quick patterns)
- Idempotency:
f(x) == f(f(x)) for normalizations/parsers
- Differential: compare two implementations or two versions, bucket on output mismatch
- Invariants: monotonic lengths, checksum equality, schema validation post‑parse
Seed Corpus
What is a Seed Corpus?
Seed corpus is a collection of input files used as a starting point for fuzzing. In mutational fuzzing, the fuzzer takes existing files and makes random mutations (flipping, reordering, removing, inserting data) before having the target application parse the mutated file.
Why is a Seed Corpus Important?
- Increases code coverage, which correlates strongly with finding crashes
- Allows fuzzers to start close to "interesting" points in the target application
- Helps overcome "unsolvable cliffs" in code coverage that fuzzers might struggle with
- Research shows fuzzers perform significantly better when bootstrapped with a minimized corpus
Creating a Seed Corpus
Manual Creation
- Create files with different command-line tools or GUI applications
- Use all possible tools for your file format to generate diverse outputs
- Automate where possible, but can be time-consuming
Existing Test Suites and Bug Reports
- Leverage test suites from the target application
- Extract test cases from bug reports (high signal value)
Web Crawling
- Use Common Crawl or similar datasets to extract relevant file types
- Filter based on MIME types
- Perform corpus distillation: keep only files that trigger new code paths
Corpus Distillation Process
- Iterate through each potential seed file
- Check file size (smaller files generally more efficient for fuzzing)
- Run the file with code coverage instrumentation
- Record if the file triggers any new code coverage not seen before
- Keep only files that increase code coverage
- Perform final minimum-set cover minimization for the smallest possible corpus
Tools for Corpus Creation
- common-corpus: Tool to build fuzzing corpus from Common Crawl data
- AFL's afl-cmin: Corpus minimization tool
- AFL's afl-tmin: Test case minimizer
- cf‑min – distributed/cluster‑friendly corpus minimizer
- Crash‑Repro‑Recorder (CRR) bundles – store the exact sequence of inputs needed to reproduce a crash deterministically
Crash Triage Pipeline
# 1) Minimize
afl-tmin -i crash -o crash.min -- ./target @@
# 2) Symbolize/Log
ASAN_OPTIONS=abort_on_error=1:symbolize=1 ./target crash.min 2>asan.log
# 3) Coverage Hash
./cov-tool --bbids ./target crash.min > cov.hash
# 4) Bucket
./bucket.py --key "$(cat cov.hash)" --log asan.log --out triage/
Cross‑platform crash analysis quick cheatsheets (user‑mode)
Linux
Windows
- Local dumps (WER):
New-Item 'HKLM:\SOFTWARE\Microsoft\Windows\Windows Error Reporting\LocalDumps' -Force | Out-Null
New-ItemProperty -Path 'HKLM:\...\LocalDumps' -Name DumpType -Type DWord -Value 2 -Force | Out-Null
- PageHeap (user mode):
gflags /p /enable target.exe /full
- WinDbg basics:
.symfix; .reload /f
!analyze -v
k; r; lm
- Optional: record Time‑Travel Debugging (TTD) to replay non‑deterministic crashes.
macOS
lldb -o 'bt all' -- ./target crash.min; set DEVELOPER_DIR to Xcode for symbols.
Kernel crash triage quicksheet
Sanitizer options (quick reference)
- ASAN_OPTIONS:
abort_on_error=1:symbolize=1:allocator_may_return_null=1:detect_stack_use_after_return=1
- UBSAN_OPTIONS:
print_stacktrace=1:halt_on_error=1
- TSAN_OPTIONS:
halt_on_error=1:history_size=7:second_deadlock_stack=1
- MSAN_OPTIONS:
poison_in_dtor=1:track_origins=2
- HWASAN (AArch64):
abort_on_error=1:stack_history_size=7
syzkaller crash repro and bisection (essentials)
# Reproduce from syz repro
syz-execprog -repeat=0 -procs=1 -wait=5m -cover=0 -debug target.repro
# Run minimized C reproducer under KASAN/KCFI build for clarity
make -j$(nproc) CONFIG_KASAN=y CONFIG_UBSAN=y CONFIG_KCFI=y
# Git bisection (when you have a fix commit)
git bisect start <bad> <good>
git bisect run ./repro.sh
Workflow
flowchart TD
A[Research Program Under Test] --> B[Choose Analysis Techniques];
B --> C[Initial Fuzzer Setup];
subgraph "Setup Details"
C --> C1[Gather Initial Seed Corpus];
C --> C2[Instrument Target];
C --> C3[Configure Execution];
C --> C4[Select Fuzzing Parameters];
end
C --> D[Begin Fuzzing];
D --> E{Monitor Fuzzing Progress};
E -- Stalled? --> F[Troubleshoot/Adjust];
F --> D;
E -- Crashes Found? --> G[Process Fuzzing Results];
subgraph "Result Processing"
G --> G1[Triage Crashes];
G --> G2[Deduplicate Test Cases];
G --> G3[Minimize Test Cases];
end
G --> H[Report Vulnerabilities / Refine];
H --> I[End];
E -- No Crashes/Finished --> I;
classDef setup fill:#99e,stroke:#111,stroke-width:2px,color:#333;
class C1,C2,C3,C4 setup;
classDef process fill:#e6e,stroke:#111,stroke-width:2px,color:#333;
class G1,G2,G3 process;
Research the Program Under Test
- Familiarize yourself with the program under test
- Learn what the program under test does and how it operates
- Interact with the program & learn how to use it
- Identify inputs & outputs
- Identify program areas to focus analysis efforts on Looking for
- Potentially Vulnerable program code
- Previously patched code
- Previous vulnerabilities
- Newly developed code
- Complex logic
- Input data ingestion sites
- For kernel modules, look beyond traditional attack vectors:
- Beyond IOCTL handlers and
copy_from_user calls
- Specialized subsystems like DMA-BUF may have attack surface through:
- Custom callbacks in operation structures
- Memory mapping handlers
- Page fault handlers
- VM operation structures
- Allocation/free mechanisms
- Identify all user-controlled inputs, especially those passed to functions that don't use
copy_from_user
Choose the Right Set of Program Analyses
- Types of analyses that you select to conduct depends on a variety of factors
- Size
- Format
- Interface
- Complexity
Initial Fuzzer Setup
Construct a representative initial seed corpus by gathering a set of program inputs that resemble the input format program expects
Instrument the program under test with coverage and sanitization
Recommended build flags (Clang/LLVM):
# LibFuzzer + ASan/UBSan (C/C++)
CC=clang CXX=clang++ CFLAGS="-O1 -g -fno-omit-frame-pointer" \
CXXFLAGS="-O1 -g -fno-omit-frame-pointer" \
LDFLAGS="" \
cmake -DCMAKE_C_FLAGS="-fsanitize=address,undefined -fsanitize-recover=undefined" \
-DCMAKE_CXX_FLAGS="-fsanitize=fuzzer,address,undefined -fsanitize-recover=undefined" ..
# MSVC (Windows) AddressSanitizer for x64
# In Visual Studio 2022+: Project Properties → C/C++ → Address Sanitizer: Yes (/fsanitize=address)
# Runtime options
set ASAN_OPTIONS=detect_leaks=1:halt_on_error=1:strict_string_checks=1
Execution method with Harness or command line arguments
Select fuzzing parameters like power schedule, dictionary, custom mutator, etc
Fuzzing setup to run in parallel, etc
Quick‑Start Recipes
LibFuzzer harness (C++):
#include <cstdint>
#include <cstddef>
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) {
/* parse_or_process(data, size); */
return 0;
}
AFL++ on a CLI target:
# Instrument
CC=afl-clang-fast CXX=afl-clang-fast++ cmake -DCMAKE_BUILD_TYPE=Release .. && make -j
# Seed dir with a few minimal valid inputs
afl-fuzz -i seeds -o findings -- ./target @@
# Useful extras: dictionary and cmplog for hard compares
afl-fuzz -i seeds -x dict.txt -o findings -c 0 -- ./target @@
# Or enable cmplog build/runner pair
AFL_LLVM_CMPLOG=1 CC=afl-clang-fast CXX=afl-clang-fast++ make clean all
afl-fuzz -i seeds -o findings -M f1 -- ./target @@
afl-fuzz -i seeds -o findings -S s1 -c 0 -- ./target @@
Windows binary‑only (QEMU mode):
afl-fuzz -Q -i seeds -o findings -- target.exe @@
Begin Fuzzing
- Start the actual fuzzing process and wait for any crashes
- Fuzzing might get stalled, become unstable or have poor performance
Process the Fuzzing Results
- Once the fuzzer produces a set of interesting test cases, we need to refine them
- Triage
- Group test cases by root cause and/or vulnerability type
- Prioritize patching the more sever vulnerabilities
- Deduplication: Remove all non-unique test cases
- Minimization: get rid of unnecessary bytes of input
Crash dedup & triage tips
- Prefer stack‑hash or coverage‑hash based bucketing (e.g., AFLTriage, Crash Triage scripts).
- Minimize before debugging:
afl-tmin, llvm-reduce, or creduce for textual formats.
- Use stable environments: pin CPU governor, disable ASLR where safe, fix random seeds.
- Export sanitizer logs to files for CI artifacts.
Reproducibility quick checklist
- Save and replay exact input sequences in persistent mode
- Pin CPU governor; fix RNG seeds; disable ASLR only where safe and necessary
- Minimize before debugging (
afl-tmin, llvm-reduce, creduce)
- Record binary hashes and sanitizer options with every crash
Obstacles
Binary Only vs Source Fuzzing
- Sometimes only executables are available, without any source code
- Solution
- Binary Rewriting: inserting coverage and sanitization into a binary without recompilation (e.g.,
retrowrite, binary‑only ASan/QASan)
- Dynamic Binary Instrumentation: inserting coverage and sanitization at runtime (e.g.,
QASAN, DynamoRIO, Frida/QBDI)
Fuzzing Harness
- Some fuzzing targets are difficult to interact with
- Solution
- Fuzzing harness acts as a middleware between the program under test and the fuzzer
- Can also use function hooking to replace network or file function calls
Library Harness Best Practices
- Research Existing Work: Search for existing harnesses or test suites
- Optimize Compilation: Use
-O1 or -O3 flags during compilation
- Balance Coverage vs Speed: A good harness maximizes code coverage while maintaining execution speed
- Strategic Targeting: Create multiple small harnesses for different library components
- Resource Management: Ensure proper initialization and cleanup of library resources
- Input Manipulation: Create functions that transform fuzzer input into valid library parameters
- Persistent Mode: For maximum efficiency, implement persistent mode harnesses that reuse library instances
Snapshot Harnessing Tips
- Snapshot after expensive initialization, right before the parse/dispatch loop
- Map fuzzer input into in‑memory buffers; avoid filesystem and network overhead
- Seed RNG and log it; ensure deterministic time sources where possible
- Export coverage early and often (breakpoints or tracepoints) for plateau detection
Structure‑aware fuzzing
- Extract tokens/keywords to a dictionary (AFL++
--dict2 or AFL_TOKEN_FILE).
- Use libprotobuf‑mutator for protobuf/gRPC targets; generate corpus from
.proto examples.
- For HTTP/REST/GraphQL, record real traffic and convert to templates with placeholders.
When building harnesses for libraries:
- Start with a basic implementation that validates core functionality
- Gradually expand to cover more API functions
- Use the documentation to understand parameter boundaries and edge cases
- Test with a variety of input encodings and configurations
- Consider structural awareness when fuzzing format-specific libraries
Resources:
Fuzzer Stalls
- Fuzzer progress has come to a halt
- Solution
- Run a collection of different fuzzers
- Produce actionable coverage statistics with tools like VisFuzz or afl-cov that inform you of where the fuzzer is getting stuck
- Use a concolic fuzzer, which combines constraints solving with fuzzing to help traditional fuzzers get past difficult conditional statements
- If all else fails, move on to a different analysis technique
- For AFL++: enable
-c 0 (cmplog), try AFL_MAP_SIZE=1048576, use -L 0 for MOpt, and add custom mutators.
Plateau escape tactics
- Switch to directed fuzzing (e.g., AFLGo/UAFuzz) for a specific function/basic block
- Add grammar/dictionary; enable CMPLOG/Redqueen to unlock hard compares
- Use concolic assistance (QSYM, Driller, libAFL concolic) on stubborn branches
- Reduce target surface via snapshotting or split harness to increase exec/s
Fuzzing Reproducibility
- Inability to reproduce crashing test cases produced by fuzzer
- Solution
- During persistent fuzzing, save all inputs starting from when a new process is created and replay those inputs in same order to reproduce behavior
- AFLPlusPlus offers special compile-time instrumentation that attempts to eliminate concurrency issues
Speed & Performance
- Fuzzer is running slowly
- Solution
- Use snapshot fuzzing to eliminate execution of redundant, unimportant code
- Avoid using complex fuzzing logic
- Parallel fuzzing can sometimes help
- Pin CPU governor to performance; disable throttling; set
AFL_SKIP_CPUFREQ=1 when needed.
- Prefer
llvm_mode over qemu_mode where possible; enable LTO instrumentation for higher throughput.
Techniques
Syzkaller
- Limit enabled syscalls to make fuzzing go deeper
- Write new
syzlang descriptions
- Change the mutation of fuzzing inputs(integrate symbolic execution)
- Start fuzzing from the crafted corpus
- Customize
kcov or use cover filter for directed fuzzing
- Extend grammar for better coverage
- External Network Fuzzing:
- Packet Injection: Utilize
TUN/TAP virtual network devices to inject network packets from within the VM, allowing the kernel to process them as if they were received externally. This approach is compatible with syzkaller's architecture, which runs the fuzzer process inside the VM.
- Coverage Collection: Employ
KCOV to gather code coverage from the kernel's network packet parsing code. KCOV can be adapted to work with TUN/TAP to trace the execution paths taken during packet processing.
- Pseudo-syscalls: Implement
syzkaller pseudo-syscalls to manage network-related operations, such as packet injection and resource management (e.g., syz_emit_ethernet for sending packets, syz_extract_tcp_res for handling TCP sequence and acknowledgement numbers).
- Syscall Descriptions: Create detailed
syzlang descriptions for network protocols and packet structures to guide the fuzzer. This includes defining packet fields, checksums, and relationships between different protocol layers.
- Integration Challenges:
- Checksums: Implement logic to correctly calculate and update checksums for various protocols (IP, TCP, UDP, ICMP) as the fuzzer mutates packet data.
- TCP Connections: Develop sequences of syscalls and pseudo-syscalls to establish and manage TCP connections, enabling fuzzing of stateful TCP communication.
- ARP Traffic: Minimize or filter out ARP traffic to isolate the fuzzing of specific protocols and avoid interference.
- IPv6 Support: Extend descriptions and logic to support IPv6, including handling extension headers and specific IPv6 features.
- Reading Code: For understanding external network fuzzing implementation, refer to the original pull request in syzkaller and the current sources, focusing on
initialize_netdevices(), syz_emit_ethernet(), and network protocol descriptions in syzlang.
- Upstream docs – See
docs/external_fuzzing_network.md in the syzkaller repo for checksum helpers and packet templates.
Target Selection
- Use syzbot coverage heatmap to identify subsystems with less-than-ideal coverage
- Look for subsystems with coverage between 1-20% (completely uncovered might have good reasons)
- Network subsystems are easier to fuzz than hardware-dependent ones
- Check syzbot dashboard to identify promising targets and current fuzzing status
Attack Surface Analysis
- Analyze kernel code to understand the attack surface (e.g., examining netlink handlers)
- Map out the available operations (e.g., socket operations, netlink commands)
- Understand the code paths that process user inputs for targeted fuzzing
Performance Optimization
- Use hardware virtualization when possible for better performance
- KVM on Linux, HVF on macOS can provide 3-5x speedup over TCG emulation
- Adapt QEMU parameters to work with available accelerators
Syzkaller Configuration and Setup
Cross-Architecture Setup:
- When configuring on non-standard hosts (e.g., ARM64 Mac), compile the target elements (kernel, rootfs) natively on Linux VM
- Go 1.23 fixes the internal linker, but you still need
CROSS_COMPILE= when the kernel uses a non‑GNU tool‑chain.
- Specify OS/arch pairs when compiling:
make HOSTOS=darwin HOSTARCH=arm64 TARGETOS=linux TARGETARCH=arm64
- Build
syz-executor on a Linux machine and copy to your host if cross-compilation fails
Kernel Module Fuzzing:
- Extract constants with
syz-extract targeting specific syscall descriptions: bin/syz-extract -os linux -sourcedir /path/to/linux -arch arm64 -build module_name.txt
- If extraction fails, manually compile programs to determine constant values
- Use the obtained values to create/fix
.const files for your module
Configuration File Example:
{
"name": "QEMU-aarch64",
"target": "linux/arm64",
"http": ":56700",
"workdir": "/path/to/workdir",
"kernel_obj": "/path/to/kernel",
"syzkaller": "/path/to/syzkaller",
"image": "/path/to/rootfs.ext3",
"sshkey": "/path/to/id_rsa",
"procs": 8,
"enable_syscalls": ["openat$module_name", "ioctl$IOCTL_CMD", "mmap"],
"type": "qemu",
"vm": {
"count": 4,
"qemu": "/path/to/qemu-system-aarch64",
"cmdline": "console=ttyAMA0 root=/dev/vda",
"kernel": "/path/to/Image",
"cpu": 2,
"mem": 2048
}
}
Kernel fuzzing practicalities
Use KCFI and KASAN builds for safer fuzzing; CONFIG_DEBUG_INFO_BTF=y helps symbolization.
Prefer TUN/TAP packet injection over raw sockets for stable replay.
Stabilize coverage with kcov filters and syz_cover_filter.
Performance Considerations:
- For ARM64 Macs, thermal throttling can significantly reduce execution rates
- Monitor execution rates for performance degradation
- VM acceleration and hardware features can improve performance
- Consider using Hardware Tag-Based KASAN for better performance vs. coverage tradeoff
- Linux 6.8+ configs
- Enable:
CONFIG_KASAN=y (or HWASAN on AArch64), CONFIG_KCSAN=y, CONFIG_UBSAN=y, CONFIG_KCFI=y, CONFIG_DEBUG_INFO_BTF=y.
- Prefer
CONFIG_KFENCE=n during fuzzing; enable later to confirm bugs.
AFL
- Crafting high quality harness
- identify existing harnesses oss-fuzz
- adapt and enhance for your fuzzing objectives
- Corpus
- use
afl-tmin and afl-cmin to tune corpus dynamically
- Code Coverage
- use
afl-cov to understand the code coverage
- then analyze and fine tune your harness
- and optimize your test corpus based on coverage feedback
- Efficient Crash Triage
Modern AFL++ tips
- Use
-M/-S for parallel fuzzer instances; combine -c 0 (cmplog) on some slaves.
- Persistent mode (
__AFL_LOOP(N)) for hot loops; in‑process fuzzing with afl++-unicorn for emulated targets.
- Dictionaries (
-x) and laf-intel/split-switches to simplify hard conditions at compile time.
MacOS
IPC Fuzzing
- We can mutate and fuzz the message that
mach_msg is sending
- Simple to do but slow
- Hard to determine which process caused the crash
- Hard to identify code coverage
- We can directly send a message to Target message handler(skipping kernel)
- Very fast and easy to instrument
- Easy to understand what caused the crash
- but Different from end exploit, might need to invoke initialization routines
- Write a Fuzzing harness
void *lib_handle = dlopen("libexample.dylib", RTLD_LAZY);
pFunction = dlsym(lib_handle, "DesiredFunction");
EDR
EDR Attack Surface Overview
EDR solutions present a significant attack surface due to their complex architecture:
Potential vulnerability categories:
- Memory corruptions in file scanning & emulation (e.g., CVE-2021-1647)
- Arbitrary file deletion via symlink vulnerabilities
- User-Mode IPC authorization issues or memory corruptions
- User-to-driver authorization bypass
- Classic driver vulnerabilities (WDM, KMDF, Mini-Filter)
- Server-to-agent authentication flaws
- Parsing bugs in server responses
- Classic Windows application vulnerabilities (DLL hijacking, file permissions)
- Emulation/sandbox escapes
- Logic bugs in security implementations
Microsoft Defender's Scanning Engine (mpengine.dll)
Microsoft Defender includes a local analysis engine mpengine.dll that performs static checks and uses emulation environments for different file types. This engine presents a significant attack surface due to its complexity and the variety of file formats it processes.
Target Characteristics:
- Installed by default on Windows systems
- Runs as SYSTEM with high privileges
- Has 1-click remote attack surface through file scanning
- Processes numerous file formats with complex parsing logic
- Prone to memory corruption vulnerabilities
Fuzzing Methodology:
Snapshot Fuzzing with WTF:
- Uses Windows Terminal Framework for coverage-guided fuzzing
- Takes snapshots after Defender initiates file scanning
- Maintains identical configuration to real environment
- Avoids limitations of manual engine bootstrapping
Alternative Approaches:
- kAFL/NYX: Additional fuzzing frameworks tested
- Jackalope: Cross-platform coverage-guided fuzzer
- Manual harness: Historic approach using
RSIG_BOOTENGINE and RSIG_SCAN_STREAMBUFFER
Practical Attack Scenarios:
Web-based DoS:
<!-- Crash Defender via malicious file download -->
<a href="crash.pdf">Download PDF</a>
<!-- Crashes MsMpEng.exe when file is scanned -->
Network Share DoS:
# Upload crash file to SMB share before credential dumping
copy crash.doc \\target\share\
# Defender crashes when scanning uploaded file
mimikatz.exe privilege::debug sekurlsa::logonpasswords
Fuzzing Setup Requirements:
Snapshot Creation:
// Example WTF harness setup
if (!g_Backend->SetBreakpoint("nt!KeBugCheck2", [](Backend_t *Backend) {
const uint64_t BCode = Backend->GetArg(0);
const std::string Filename = fmt::format("crash-{:#x}", BCode);
Backend->Stop(Crash_t(Filename));
}))
Coverage Analysis:
- Use IDA Lighthouse for visualization
- Monitor for DRIVER_VERIFIER_DETECTED_VIOLATION (0xc4)
- Track IRQL_NOT_LESS_OR_EQUAL (0xa) crashes
Defense Implications:
- Demonstrates need for robust input validation
- Shows importance of crash recovery mechanisms
- Highlights risks of complex file parsing engines
- Suggests value of sandboxing scanning processes
Cross-platform mpengine.dll Fuzzing
Recent advances enable fuzzing the latest Windows Defender engine (v1.1.25020.1007) on Linux using loadlibrary with Intel PT coverage:
// Disable Lua VM to avoid stability issues
void my_lua_exec(){
return;
}
int main(int argc, char** argv){
// Hook luaV_execute to bypass Lua signature processing
insert_function_redirect((void*)luaV_execute_address, my_lua_exec, HOOK_REPLACE_FUNCTION);
// Setup persistent fuzzing loop
for (;;) {
size_t len;
uint8_t *buf;
HF_ITER(&buf, &len);
ScanDescriptor.UserPtr = fmemopen(buf, len, "r");
if (__rsignal(&KernelHandle, RSIG_SCAN_STREAMBUFFER, &ScanParams, sizeof ScanParams) != 0) {
// Handle scan results
}
}
}
Performance Optimizations
- Persistent Mode: Eliminates initialization overhead, achieving hundreds of exec/s
- Intel PT Coverage: Hardware-based tracing for binary-only targets
- Lua VM Bypassing: Reduces crashes and focuses on native code vulnerabilities
honggfuzz with Intel PT Setup
# Non-persistent mode (slower, ~4s/execution)
../honggfuzz/honggfuzz -i ~/input/ -W ~/workspace/ --linux_perf_ipt_block -t 10 -- ./mpclient_x64 ___FILE___
# Per
…(truncated)
1---2name: offensive-fuzzing3description: SKILL: Fuzzing4---5# SKILL: Fuzzing67## Metadata8- **Skill Name**: fuzzing-methodology9- **Folder**: offensive-fuzzing10- **Source**: https://github.com/SnailSploit/offensive-checklist/blob/main/fuzzing.md1112## Description13Practical fuzzing methodology reference: target identification, fuzzer selection (AFL++, libFuzzer, Boofuzz, Peach), harness writing, corpus curation, mutation strategies, coverage measurement, and crash triage. Use when setting up or running fuzz campaigns against any target.1415## Trigger Phrases16Use this skill when the conversation involves any of:17`fuzzing, AFL++, libFuzzer, Boofuzz, Peach, fuzz harness, corpus, mutation, coverage, crash triage, network fuzzing, file format fuzzing`1819## Instructions for Claude2021When this skill is active:221. Load and apply the full methodology below as your operational checklist232. Follow steps in order unless the user specifies otherwise243. For each technique, consider applicability to the current target/context254. Track which checklist items have been completed265. Suggest next steps based on findings2728---2930## Full Methodology3132# Fuzzing3334automated software testing technique that is used for program analysis3536## Types3738### BlackBox3940- Generates random or model‑based inputs and feeds them into the program under test41- Fuzzer has no knowledge of the program internals during fuzzing42- Pros43 - Extremely fast44 - Easy to use45 - Scalable46- Cons47 - Random testing leads to poor coverage4849### GreyBox5051- Relies on lightweight instrumentation of the program under test52- Fuzzer has some knowledge of the program internals during fuzzing53- Pros54 - Feedback55 - Scalable56 - Relatively fast57- Cons58 - Simplistic mutations59- Examples60 - General: [AFL++](https://github.com/AFLplusplus/AFLplusplus), [Honggfuzz](https://github.com/google/honggfuzz), [BooFuzz](https://github.com/jtpereyda/boofuzz), [WinAFL](https://github.com/googleprojectzero/winafl)61 - Directed: [UAFuzz](https://github.com/strongcourage/uafuzz), [AFLGo](https://github.com/aflgo/aflgo)62 - Grammar Based: [Tlspuffin](https://github.com/tlspuffin/tlspuffin), [AFLSmart](https://github.com/aflsmart/aflsmart)63 - Taint Based: [Angora](https://github.com/AngoraFuzzer/Angora),64 - Concolic: [QSYM](https://github.com/sslab-gatech/qsym), [AFL](https://aflplus.plus/libafl-book/advanced_features/concolic/concolic.html), [SymSan](https://github.com/R-Fuzz/symsan)65 - Self‑learning: [SieveFuzz](https://github.com/SieveFuzz/SieveFuzz), [RL‑Fuzz](https://github.com/fuzzwareorg/rl-fuzz)66 - Binary Fuzzing: [Jackalope](https://github.com/googleprojectzero/Jackalope)67 - Special Fuzzers: [Papora](https://github.com/ambergroup-labs/papora), [Medusa](https://github.com/trailofbits/medusa)68 - Kernel: [syzkaller](https://github.com/google/syzkaller), [kAFL](https://github.com/IntelLabs/kAFL), [wtf](https://github.com/0vercl0k/wtf), [KF/x](https://github.com/intel/kernel-fuzzer-for-xen-project)69 - ETC: [LibAFL](https://github.com/AFLplusplus/LibAFL) (snapshot & Unicorn support 0.15.x), [FuzzBench](https://google.github.io/fuzzbench/), [ClusterFuzz](https://github.com/google/clusterfuzz), [DynamoRIO](https://github.com/DynamoRIO/dynamorio)7071### Snapshot7273- Takes and restores memory/register snapshots of the target or VM to bypass expensive initialization.74- Pros75 - Order‑of‑magnitude faster execution on large or stateful binaries76 - Deterministic and easily reproducible77 - Works with closed‑source targets (no rebuild needed)78- Cons79 - Initial snapshot & harness creation can be complex80 - Requires specialised snapshot engine support (e.g., Nyx, Snapchange, wtf)81- Examples82 - [Nyx](https://github.com/nyx-fuzz/Nyx) (v2 with Intel‑PT support; note: Nyx integration requires the AFL++-Nyx fork, not mainline AFL++ 4.x)83 - [Snapchange](https://github.com/awslabs/snapchange)84 - [wtf](https://github.com/0vercl0k/wtf)8586#### Snapshot Fuzzing Recipes8788- **AFL++ Nyx (user-mode)**8990```bash91NYX_MODE=1 AFL_MAP_SIZE=1048576 afl-fuzz -i seeds -o findings -- ./target_nyx @@92```9394### WhiteBox9596- Leverages heavyweight symbolic execution and static analysis, assuming full source availability97- Pros98 - Solves path constraints to get past difficult logic99 - Capable of generating inputs for deep, hard‑to‑reach program states100- Cons101 - Complex102 - Slow103 - Not scalable104105### Ensemble Fuzzing106107Ensemble fuzzing coordinates multiple heterogeneous fuzzers sharing a unified corpus, achieving better coverage than any single fuzzer.108109#### Concept110111- **Cross-Pollination:** Different fuzzers discover different code paths; sharing corpus maximizes coverage112- **Complementary Strengths:** AFL++ excels at control flow; Honggfuzz at crash detection; libFuzzer at in-process fuzzing113- **Coordinated Scheduling:** Use reinforcement learning or heuristics to select optimal fuzzer per input114115#### Tools and Frameworks116117- **EnFuzz:** Uses reinforcement learning to dynamically select which fuzzer runs on each input based on historical effectiveness118- **CollAFL:** Lightweight synchronization protocol for parallel fuzzer coordination without central coordinator119- **CUPID:** Analyzes feedback from all fuzzers to guide mutation strategies across the ensemble120121#### Practical Setup122123```bash124# Terminal 1: AFL++ primary with shared sync directory125afl-fuzz -M fuzzer1 -i seeds -o sync_dir -- ./target @@126127# Terminal 2: Honggfuzz syncing to AFL corpus128../honggfuzz/honggfuzz -i sync_dir/fuzzer1/queue -W sync_dir/hfuzz \129 --linux_perf_ipt_block -t 10 -- ./target ___FILE___130131# Terminal 3: libFuzzer reading from both corpora132./target_libfuzzer sync_dir/fuzzer1/queue sync_dir/hfuzz/queue \133 -max_total_time=3600 -runs=1000000134135# Terminal 4: Sync monitor (optional)136watch -n 60 'ls -lh sync_dir/*/queue | wc -l'137```138139#### Best Practices140141- **Start Simple:** Begin with AFL++ + Honggfuzz; add libFuzzer if you have source142- **Corpus Deduplication:** Run `afl-cmin` periodically to minimize combined corpus143- **Resource Allocation:** Allocate 40% AFL++, 30% Honggfuzz, 30% libFuzzer based on empirical results144- **Monitoring:** Track unique crashes per fuzzer to identify which contributes most145146#### Performance Gains147148- Research shows **15-40% more coverage** vs single fuzzer149- **2-3× more unique crashes** in 24-hour runs150- Particularly effective on complex parsers with multiple code paths151152## Components153154### Power Scheduler155156- Responsible for distributing the amount of fuzzing time among the seeds in the seed queue157- Prioritize more interesting seeds158 - Usually measure by its ability to produce new code coverage159- Assigns a scope to each seed and picks the one with highest score160- Self‑learning schedulers such as _SieveFuzz_ and _RL‑Fuzz_ use reinforcement learning to allocate energy dynamically based on past mutation success.161- Explore/Exploit auto‑tuning: AFL++ 4.30+ folds MOpt into the core and dynamically balances energy between seeds.162163### Mutation164165- Responsible for making small changes to the fuzzing seed, with the goal of exercising new program behavior166- Examples: bit-flips, addition/subtraction, deletion, etc167- Typically operate by making small, incremental changes168- LLM‑assisted mutations & harness scaffolding: Large‑language models can infer grammars/dictionaries (`--dict2`) and draft starter harness code (e.g., LibAFL's `llm_mutator`).169- Overflow‑aware arithmetic heuristics: Enable `--arith8/16/32-smart` in AFL++ (or equivalents) to target boundary‑condition overflows.170- AFLPlusPlus171 - Deterministic: includes single deterministic mutations on the content of the test cases172 - Havoc: mutations are randomly stacked and also includes changes to the size of the test case173 - Note: honggfuzz and libFuzzer implement weighted random mutation stacks instead of the deterministic/havoc split174- cmp_log guided mutations: Enable `-cmp_log` (AFL++) or `AFL_LLVM_CMPLOG=1` to perform Redqueen‑style input‑to‑state transforms.175- Snapshot‑aware mutations: With Nyx mode (`NYX_MODE=1`) AFL++ randomises in‑memory state as well as file input for deep state machines.176- LLM‑guided edits: _HyLLfuzz_ uses GPT/Llama‑3 to generate branch‑targeted mutations, achieving ≈ 1.3× edge coverage.177178### Directed Fuzzing (reach a specific bug/function)179180- AFLGo basics:181 - Build with distance metrics to target BBs/functions, then run with exploration/exploitation phases182 - Example:183 ```bash184 export AFLGO=/path/to/aflgo185 cmake -DAFLGO=ON -DDISTANCE_CFG=BBtargets.txt ..186 make clean all187 afl-fuzz -z exp -c 45m -i seeds -o out -- ./target @@188 ```189- libAFL targeted stages: use objective feedbacks to push towards addresses/symbols of interest190191### Executor192193- Responsible for executing the program under test with mutated fuzzing seed in a performant manner194- AFL++ uses a forkserver to skip the cost of expensive initialization and startup code195- AFL++ also offers persistent fuzzing mode, where the essential code is run inside a loop without a need for fork196- GPU‑accelerated execution loops are possible with _CUDA‑AFL_ and LibAFL's `VulkanExecutor`, ideal for shader or graphics targets.197- Recent AFL++ versions ship a Nyx executor capable of high exec/s on user‑mode targets; activate with `NYX_MODE=1` (verify version and build flags).198199#### Persistent Mode Implementation200201**HF_ITER Style (Honggfuzz):**202203Honggfuzz offers HF_ITER-style persistent mode that's often easier to implement than ASAN-style for complex targets:204205```cpp206#include "honggfuzz.h"207208int main(int argc, char** argv) {209 // Expensive initialization (run once)210 initialize_target();211212 // Persistent fuzzing loop213 for (;;) {214 size_t len;215 uint8_t *buf;216217 // Get input from honggfuzz218 HF_ITER(&buf, &len);219220 // Convert to target-appropriate format221 FILE* input_stream = fmemopen(buf, len, "r");222223 // Execute target with fuzzed input224 int result = target_function(input_stream);225226 // Cleanup for next iteration227 fclose(input_stream);228 reset_target_state();229 }230}231```232233**Memory Management Considerations:**234235```cpp236// Prevent memory leaks in persistent mode237void reset_target_state() {238 // Free any allocated memory239 cleanup_buffers();240241 // Reset global state242 memset(&global_state, 0, sizeof(global_state));243244 // Close any open file handles245 close_handles();246}247```248249**ASAN-Style Alternative:**250251```cpp252extern "C" int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) {253 // Direct fuzzing interface - no manual loop needed254 return target_function(data, size);255}256```257258### Feedback259260- Set of program state information from running with a particular input seed261- Fuzzers use feedback to determine fitness of a given seed(basic block coverage, edge coverage, path coverage)262- AFLPlusPlus uses edge coverage (keeps a shared bitmap between the target and the fuzzer)263- AFL++ uses edge coverage (keeps a shared bitmap between the target and the fuzzer)264- Hardware trace sources such as Intel® Processor Trace (IPT) and ARM ETM/CoreSight provide near‑zero‑overhead edge coverage for large binaries.265- LibAFL 0.15.2 can derive edge coverage from Last Branch Records, giving zero‑instrumentation tracing on recent Intel CPUs (`LBRFeedback`).266267#### Intel PT Coverage Setup268269**Honggfuzz with Intel PT:**270271Intel PT provides hardware-based coverage tracking ideal for binary-only targets where source instrumentation isn't possible:272273```bash274# Basic Intel PT setup with honggfuzz275../honggfuzz/honggfuzz -i input_corpus/ -W workspace/ --linux_perf_ipt_block -t 10 -- ./target @@276277# Performance tuning278../honggfuzz/honggfuzz -i input_corpus/ -W workspace/ \279 --linux_perf_ipt_block \280 --linux_perf_ignore_unknown \281 --threads 4 \282 -t 30 -- ./target @@283```284285**Intel PT Requirements:**286287- Modern Intel CPU with PT support (check with `cat /proc/cpuinfo | grep intel_pt`)288- Linux kernel with PT enabled (`CONFIG_INTEL_PT=y`)289- Sufficient privileges or `CAP_SYS_ADMIN` capability290291**Troubleshooting Intel PT:**292293```bash294# Check PT availability295dmesg | grep -i "intel.*pt"296cat /sys/devices/intel_pt/type297298# Permission issues299echo 'kernel.perf_event_paranoid = -1' >> /etc/sysctl.conf300sysctl -p301302# Alternative: use capsh for specific capability303capsh --caps="cap_sys_admin+eip" -- -c "./honggfuzz_command"304```305306### Oracle307308- Detects if any interesting behavior occurred during execution of the program under test with a given input309- Code Sanitizers are special compile-time instrumentation for the program under test that perform run-time checks310- AFL++ offers `QASAN` to run binaries that are not instrumented with `ASan` under QEMU with the AFL++ instrumentation311- Different Bugs Require Different Oracles312 - **No Source**: Statically rewrite the binary or `YOLO` it313 - **Decoupled Sanitization**: use [SAND](https://github.com/SANDTeam/sand) to off-load sanitizer checks and speed up fuzzing314 - **Memory Safety**: use Address or Leak or Memory Sanitizer (try HWASAN for lower memory usage); or use ARM's MTE for hardware‑assisted tagging on AArch64315 - **Concurrency Issues**: use Thread Sanitizer – LLVM 18 adds detection for C++20 std::atomic_wait316 - **Type Safety**: use Type Sanitizer317 - **Undefined Behavior**: use `UBSan`318 - **Heap Hardening**: Scudo Hardened Allocator (part of LLVM's compiler-rt, available in recent Clang versions) detects overflows and double‑frees with minimal performance cost.319 - **Kernel Undefined Behavior**: Kernel UBSan (KUBSan) — enable with `CONFIG_UBSAN_TRAP=y` (verify availability in your kernel).320 - **Control‑Flow Integrity**: **KCFI** (Clang 18) adds low-overhead forward-edge CFI; enable with `-fsanitize=kcfi`.321 - **Data‑Flow**: Data‑Flow Sanitizer v2 adds taint‑tracking (`-fsanitize=dataflow`) and now works with libFuzzer.322 - **Security‑property oracles**: Idempotency or differential fuzzing for APIs (e.g., _DiffSink_ for compiler targets).323 - **Kernel Stuff**: use `KASAN`, `KMSAN`, `KCSAN`324 - **Logic Bugs**: Differential or Property(Correctness-Idempotency) Oracles325326#### Property/Differential Oracles (quick patterns)327328- Idempotency: `f(x) == f(f(x))` for normalizations/parsers329- Differential: compare two implementations or two versions, bucket on output mismatch330- Invariants: monotonic lengths, checksum equality, schema validation post‑parse331332## Seed Corpus333334### What is a Seed Corpus?335336Seed corpus is a collection of input files used as a starting point for fuzzing. In mutational fuzzing, the fuzzer takes existing files and makes random mutations (flipping, reordering, removing, inserting data) before having the target application parse the mutated file.337338### Why is a Seed Corpus Important?339340- Increases code coverage, which correlates strongly with finding crashes341- Allows fuzzers to start close to "interesting" points in the target application342- Helps overcome "unsolvable cliffs" in code coverage that fuzzers might struggle with343- Research shows fuzzers perform significantly better when bootstrapped with a minimized corpus344345### Creating a Seed Corpus3463471. **Manual Creation**348 - Create files with different command-line tools or GUI applications349 - Use all possible tools for your file format to generate diverse outputs350 - Automate where possible, but can be time-consuming3513522. **Existing Test Suites and Bug Reports**353 - Leverage test suites from the target application354 - Extract test cases from bug reports (high signal value)3553563. **Web Crawling**357 - Use Common Crawl or similar datasets to extract relevant file types358 - Filter based on MIME types359 - Perform corpus distillation: keep only files that trigger new code paths360361### Corpus Distillation Process3623631. Iterate through each potential seed file3642. Check file size (smaller files generally more efficient for fuzzing)3653. Run the file with code coverage instrumentation3664. Record if the file triggers any new code coverage not seen before3675. Keep only files that increase code coverage3686. Perform final minimum-set cover minimization for the smallest possible corpus369370### Tools for Corpus Creation371372- [common-corpus](https://github.com/benhawkes/common-corpus): Tool to build fuzzing corpus from Common Crawl data373- [AFL's afl-cmin](https://github.com/AFLplusplus/AFLplusplus/blob/stable/afl-cmin): Corpus minimization tool374- [AFL's afl-tmin](https://github.com/AFLplusplus/AFLplusplus/blob/stable/afl-tmin): Test case minimizer375- [cf‑min](https://github.com/AFLplusplus/afl-cmin) – distributed/cluster‑friendly corpus minimizer376- Crash‑Repro‑Recorder (CRR) bundles – store the exact sequence of inputs needed to reproduce a crash deterministically377378### Crash Triage Pipeline379380```bash381# 1) Minimize382afl-tmin -i crash -o crash.min -- ./target @@383# 2) Symbolize/Log384ASAN_OPTIONS=abort_on_error=1:symbolize=1 ./target crash.min 2>asan.log385# 3) Coverage Hash386./cov-tool --bbids ./target crash.min > cov.hash387# 4) Bucket388./bucket.py --key "$(cat cov.hash)" --log asan.log --out triage/389```390391#### Cross‑platform crash analysis quick cheatsheets (user‑mode)392393- Linux394 - Enable coredumps and replay:395 ```bash396 ulimit -c unlimited397 sysctl -w kernel.core_pattern=core.%e.%p398 ./target crash.min # produces core399 gdb -q ./target core.* -ex 'set pagination off' -ex bt -ex 'info reg' -ex q400 addr2line -e ./target 0xDEADBEAF401 ```402 - Stabilize runs: `taskset -c 0 chrt -f 99 ./target @@`; set `ASAN_OPTIONS=allocator_may_return_null=1:handle_abort=1`.403404- Windows405 - Local dumps (WER):406 ```powershell407 New-Item 'HKLM:\SOFTWARE\Microsoft\Windows\Windows Error Reporting\LocalDumps' -Force | Out-Null408 New-ItemProperty -Path 'HKLM:\...\LocalDumps' -Name DumpType -Type DWord -Value 2 -Force | Out-Null409 ```410 - PageHeap (user mode):411 ```cmd412 gflags /p /enable target.exe /full413 ```414 - WinDbg basics:415 ```text416 .symfix; .reload /f417 !analyze -v418 k; r; lm419 ```420 - Optional: record Time‑Travel Debugging (TTD) to replay non‑deterministic crashes.421422- macOS423 - `lldb -o 'bt all' -- ./target crash.min`; set `DEVELOPER_DIR` to Xcode for symbols.424425#### Kernel crash triage quicksheet426427- Linux428 - KASAN/KMSAN logs: `dmesg -T | egrep -i 'kasan|kmsan' -A 60`429 - Decode stacks:430 ```bash431 ./scripts/decode_stacktrace.sh vmlinux /lib/modules/$(uname -r)/build < dmesg.log432 ```433 - Map addresses: `addr2line -e vmlinux 0xffffffff81234567`434435- Windows436 - Driver Verifier:437 ```cmd438 verifier /standard /driver yourdrv.sys439 ```440 - KD/WinDbg:441 ```text442 !analyze -v; kv; r; !verifier 3; !irpfind443 ```444445#### Sanitizer options (quick reference)446447- ASAN_OPTIONS: `abort_on_error=1:symbolize=1:allocator_may_return_null=1:detect_stack_use_after_return=1`448- UBSAN_OPTIONS: `print_stacktrace=1:halt_on_error=1`449- TSAN_OPTIONS: `halt_on_error=1:history_size=7:second_deadlock_stack=1`450- MSAN_OPTIONS: `poison_in_dtor=1:track_origins=2`451- HWASAN (AArch64): `abort_on_error=1:stack_history_size=7`452453#### syzkaller crash repro and bisection (essentials)454455```bash456# Reproduce from syz repro457syz-execprog -repeat=0 -procs=1 -wait=5m -cover=0 -debug target.repro458459# Run minimized C reproducer under KASAN/KCFI build for clarity460make -j$(nproc) CONFIG_KASAN=y CONFIG_UBSAN=y CONFIG_KCFI=y461462# Git bisection (when you have a fix commit)463git bisect start <bad> <good>464git bisect run ./repro.sh465```466467## Workflow468469```mermaid470flowchart TD471 A[Research Program Under Test] --> B[Choose Analysis Techniques];472 B --> C[Initial Fuzzer Setup];473 subgraph "Setup Details"474 C --> C1[Gather Initial Seed Corpus];475 C --> C2[Instrument Target];476 C --> C3[Configure Execution];477 C --> C4[Select Fuzzing Parameters];478 end479 C --> D[Begin Fuzzing];480 D --> E{Monitor Fuzzing Progress};481 E -- Stalled? --> F[Troubleshoot/Adjust];482 F --> D;483 E -- Crashes Found? --> G[Process Fuzzing Results];484 subgraph "Result Processing"485 G --> G1[Triage Crashes];486 G --> G2[Deduplicate Test Cases];487 G --> G3[Minimize Test Cases];488 end489 G --> H[Report Vulnerabilities / Refine];490 H --> I[End];491 E -- No Crashes/Finished --> I;492493 classDef setup fill:#99e,stroke:#111,stroke-width:2px,color:#333;494 class C1,C2,C3,C4 setup;495 classDef process fill:#e6e,stroke:#111,stroke-width:2px,color:#333;496 class G1,G2,G3 process;497```498499### Research the Program Under Test500501- Familiarize yourself with the program under test502- Learn what the program under test does and how it operates503- Interact with the program & learn how to use it504- Identify inputs & outputs505- Identify program areas to focus analysis efforts on Looking for506 - Potentially Vulnerable program code507 - Previously patched code508 - Previous vulnerabilities509 - Newly developed code510 - Complex logic511 - Input data ingestion sites512- For kernel modules, look beyond traditional attack vectors:513 - Beyond IOCTL handlers and `copy_from_user` calls514 - Specialized subsystems like DMA-BUF may have attack surface through:515 - Custom callbacks in operation structures516 - Memory mapping handlers517 - Page fault handlers518 - VM operation structures519 - Allocation/free mechanisms520 - Identify all user-controlled inputs, especially those passed to functions that don't use `copy_from_user`521522### Choose the Right Set of Program Analyses523524- Types of analyses that you select to conduct depends on a variety of factors525 - Size526 - Format527 - Interface528 - Complexity529530### Initial Fuzzer Setup531532- Construct a representative initial seed corpus by gathering a set of program inputs that resemble the input format program expects533- Instrument the program under test with coverage and sanitization534- Recommended build flags (Clang/LLVM):535536 ```bash537 # LibFuzzer + ASan/UBSan (C/C++)538 CC=clang CXX=clang++ CFLAGS="-O1 -g -fno-omit-frame-pointer" \539 CXXFLAGS="-O1 -g -fno-omit-frame-pointer" \540 LDFLAGS="" \541 cmake -DCMAKE_C_FLAGS="-fsanitize=address,undefined -fsanitize-recover=undefined" \542 -DCMAKE_CXX_FLAGS="-fsanitize=fuzzer,address,undefined -fsanitize-recover=undefined" ..543544 # MSVC (Windows) AddressSanitizer for x64545 # In Visual Studio 2022+: Project Properties → C/C++ → Address Sanitizer: Yes (/fsanitize=address)546 # Runtime options547 set ASAN_OPTIONS=detect_leaks=1:halt_on_error=1:strict_string_checks=1548 ```549550- Execution method with Harness or command line arguments551- Select fuzzing parameters like power schedule, dictionary, custom mutator, etc552- Fuzzing setup to run in parallel, etc553554### Quick‑Start Recipes555556- LibFuzzer harness (C++):557558 ```cpp559 #include <cstdint>560 #include <cstddef>561 extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) {562 /* parse_or_process(data, size); */563 return 0;564 }565 ```566567- AFL++ on a CLI target:568569 ```bash570 # Instrument571 CC=afl-clang-fast CXX=afl-clang-fast++ cmake -DCMAKE_BUILD_TYPE=Release .. && make -j572 # Seed dir with a few minimal valid inputs573 afl-fuzz -i seeds -o findings -- ./target @@574 # Useful extras: dictionary and cmplog for hard compares575 afl-fuzz -i seeds -x dict.txt -o findings -c 0 -- ./target @@576 # Or enable cmplog build/runner pair577 AFL_LLVM_CMPLOG=1 CC=afl-clang-fast CXX=afl-clang-fast++ make clean all578 afl-fuzz -i seeds -o findings -M f1 -- ./target @@579 afl-fuzz -i seeds -o findings -S s1 -c 0 -- ./target @@580 ```581582- Windows binary‑only (QEMU mode):583584 ```bash585 afl-fuzz -Q -i seeds -o findings -- target.exe @@586 ```587588### Begin Fuzzing589590- Start the actual fuzzing process and wait for any crashes591- Fuzzing might get stalled, become unstable or have poor performance592593### Process the Fuzzing Results594595- Once the fuzzer produces a set of interesting test cases, we need to refine them596- Triage597 - Group test cases by root cause and/or vulnerability type598 - Prioritize patching the more sever vulnerabilities599- Deduplication: Remove all non-unique test cases600- Minimization: get rid of unnecessary bytes of input601602#### Crash dedup & triage tips603604- Prefer stack‑hash or coverage‑hash based bucketing (e.g., AFLTriage, Crash Triage scripts).605- Minimize before debugging: `afl-tmin`, `llvm-reduce`, or `creduce` for textual formats.606- Use stable environments: pin CPU governor, disable ASLR where safe, fix random seeds.607- Export sanitizer logs to files for CI artifacts.608609#### Reproducibility quick checklist610611- Save and replay exact input sequences in persistent mode612- Pin CPU governor; fix RNG seeds; disable ASLR only where safe and necessary613- Minimize before debugging (`afl-tmin`, `llvm-reduce`, `creduce`)614- Record binary hashes and sanitizer options with every crash615616## Obstacles617618### Binary Only vs Source Fuzzing619620- Sometimes only executables are available, without any source code621- Solution622 - **Binary Rewriting**: inserting coverage and sanitization into a binary without recompilation (e.g., `retrowrite`, binary‑only ASan/QASan)623 - **Dynamic Binary Instrumentation**: inserting coverage and sanitization at runtime (e.g., `QASAN`, DynamoRIO, Frida/QBDI)624625### Fuzzing Harness626627- Some fuzzing targets are difficult to interact with628- Solution629 - Fuzzing harness acts as a middleware between the program under test and the fuzzer630 - Can also use function hooking to replace network or file function calls631632#### Library Harness Best Practices633634- **Research Existing Work**: Search for existing harnesses or test suites635- **Optimize Compilation**: Use `-O1` or `-O3` flags during compilation636- **Balance Coverage vs Speed**: A good harness maximizes code coverage while maintaining execution speed637- **Strategic Targeting**: Create multiple small harnesses for different library components638- **Resource Management**: Ensure proper initialization and cleanup of library resources639- **Input Manipulation**: Create functions that transform fuzzer input into valid library parameters640- **Persistent Mode**: For maximum efficiency, implement persistent mode harnesses that reuse library instances641642#### Snapshot Harnessing Tips643644- Snapshot after expensive initialization, right before the parse/dispatch loop645- Map fuzzer input into in‑memory buffers; avoid filesystem and network overhead646- Seed RNG and log it; ensure deterministic time sources where possible647- Export coverage early and often (breakpoints or tracepoints) for plateau detection648649#### Structure‑aware fuzzing650651- Extract tokens/keywords to a dictionary (AFL++ `--dict2` or `AFL_TOKEN_FILE`).652- Use libprotobuf‑mutator for protobuf/gRPC targets; generate corpus from `.proto` examples.653- For HTTP/REST/GraphQL, record real traffic and convert to templates with placeholders.654655When building harnesses for libraries:6566571. Start with a basic implementation that validates core functionality6582. Gradually expand to cover more API functions6593. Use the documentation to understand parameter boundaries and edge cases6604. Test with a variety of input encodings and configurations6615. Consider structural awareness when fuzzing format-specific libraries662663Resources:664665- [Awesome LibFuzzer Harness Collection](https://github.com/google/fuzzing/tree/master/examples)666- OSS-Fuzz project repositories667668### Fuzzer Stalls669670- Fuzzer progress has come to a halt671- Solution672 - Run a collection of different fuzzers673 - Produce actionable coverage statistics with tools like VisFuzz or afl-cov that inform you of where the fuzzer is getting stuck674 - Use a concolic fuzzer, which combines constraints solving with fuzzing to help traditional fuzzers get past difficult conditional statements675 - If all else fails, move on to a different analysis technique676 - For AFL++: enable `-c 0` (cmplog), try `AFL_MAP_SIZE=1048576`, use `-L 0` for MOpt, and add custom mutators.677678#### Plateau escape tactics679680- Switch to directed fuzzing (e.g., AFLGo/UAFuzz) for a specific function/basic block681- Add grammar/dictionary; enable CMPLOG/Redqueen to unlock hard compares682- Use concolic assistance (QSYM, Driller, libAFL concolic) on stubborn branches683- Reduce target surface via snapshotting or split harness to increase exec/s684685### Fuzzing Reproducibility686687- Inability to reproduce crashing test cases produced by fuzzer688- Solution689 - During persistent fuzzing, save all inputs starting from when a new process is created and replay those inputs in same order to reproduce behavior690 - AFLPlusPlus offers special compile-time instrumentation that attempts to eliminate concurrency issues691692### Speed & Performance693694- Fuzzer is running slowly695- Solution696 - Use snapshot fuzzing to eliminate execution of redundant, unimportant code697 - Avoid using complex fuzzing logic698 - Parallel fuzzing can sometimes help699 - Pin CPU governor to performance; disable throttling; set `AFL_SKIP_CPUFREQ=1` when needed.700 - Prefer `llvm_mode` over `qemu_mode` where possible; enable LTO instrumentation for higher throughput.701702## Techniques703704### Syzkaller705706- Limit enabled syscalls to make fuzzing go deeper707- Write new `syzlang` descriptions708- Change the mutation of fuzzing inputs(integrate symbolic execution)709- Start fuzzing from the crafted corpus710- Customize `kcov` or use cover filter for directed fuzzing711- Extend grammar for better coverage712- **External Network Fuzzing**:713 - **Packet Injection**: Utilize `TUN/TAP` virtual network devices to inject network packets from within the VM, allowing the kernel to process them as if they were received externally. This approach is compatible with syzkaller's architecture, which runs the fuzzer process inside the VM.714 - **Coverage Collection**: Employ `KCOV` to gather code coverage from the kernel's network packet parsing code. `KCOV` can be adapted to work with `TUN/TAP` to trace the execution paths taken during packet processing.715 - **Pseudo-syscalls**: Implement `syzkaller` pseudo-syscalls to manage network-related operations, such as packet injection and resource management (e.g., `syz_emit_ethernet` for sending packets, `syz_extract_tcp_res` for handling TCP sequence and acknowledgement numbers).716 - **Syscall Descriptions**: Create detailed `syzlang` descriptions for network protocols and packet structures to guide the fuzzer. This includes defining packet fields, checksums, and relationships between different protocol layers.717 - **Integration Challenges**:718 - **Checksums**: Implement logic to correctly calculate and update checksums for various protocols (IP, TCP, UDP, ICMP) as the fuzzer mutates packet data.719 - **TCP Connections**: Develop sequences of syscalls and pseudo-syscalls to establish and manage TCP connections, enabling fuzzing of stateful TCP communication.720 - **ARP Traffic**: Minimize or filter out ARP traffic to isolate the fuzzing of specific protocols and avoid interference.721 - **IPv6 Support**: Extend descriptions and logic to support IPv6, including handling extension headers and specific IPv6 features.722 - **Reading Code**: For understanding external network fuzzing implementation, refer to the original pull request in syzkaller and the current sources, focusing on `initialize_netdevices()`, `syz_emit_ethernet()`, and network protocol descriptions in `syzlang`.723 - **Upstream docs** – See `docs/external_fuzzing_network.md` in the syzkaller repo for checksum helpers and packet templates.724725#### Target Selection726727- Use syzbot coverage heatmap to identify subsystems with less-than-ideal coverage728- Look for subsystems with coverage between 1-20% (completely uncovered might have good reasons)729- Network subsystems are easier to fuzz than hardware-dependent ones730- Check syzbot dashboard to identify promising targets and current fuzzing status731732#### Attack Surface Analysis733734- Analyze kernel code to understand the attack surface (e.g., examining netlink handlers)735- Map out the available operations (e.g., socket operations, netlink commands)736- Understand the code paths that process user inputs for targeted fuzzing737738#### Performance Optimization739740- Use hardware virtualization when possible for better performance741 - KVM on Linux, HVF on macOS can provide 3-5x speedup over TCG emulation742- Adapt QEMU parameters to work with available accelerators743744#### Syzkaller Configuration and Setup745746- **Cross-Architecture Setup**:747 - When configuring on non-standard hosts (e.g., ARM64 Mac), compile the target elements (kernel, rootfs) natively on Linux VM748 - Go 1.23 fixes the internal linker, but you **still need `CROSS_COMPILE=`** when the kernel uses a non‑GNU tool‑chain.749 - Specify OS/arch pairs when compiling: `make HOSTOS=darwin HOSTARCH=arm64 TARGETOS=linux TARGETARCH=arm64`750 - Build `syz-executor` on a Linux machine and copy to your host if cross-compilation fails751752- **Kernel Module Fuzzing**:753 - Extract constants with `syz-extract` targeting specific syscall descriptions: `bin/syz-extract -os linux -sourcedir /path/to/linux -arch arm64 -build module_name.txt`754 - If extraction fails, manually compile programs to determine constant values755 - Use the obtained values to create/fix `.const` files for your module756757- **Configuration File Example**:758759 ```json760 {761 "name": "QEMU-aarch64",762 "target": "linux/arm64",763 "http": ":56700",764 "workdir": "/path/to/workdir",765 "kernel_obj": "/path/to/kernel",766 "syzkaller": "/path/to/syzkaller",767 "image": "/path/to/rootfs.ext3",768 "sshkey": "/path/to/id_rsa",769 "procs": 8,770 "enable_syscalls": ["openat$module_name", "ioctl$IOCTL_CMD", "mmap"],771 "type": "qemu",772 "vm": {773 "count": 4,774 "qemu": "/path/to/qemu-system-aarch64",775 "cmdline": "console=ttyAMA0 root=/dev/vda",776 "kernel": "/path/to/Image",777 "cpu": 2,778 "mem": 2048779 }780 }781 ```782783#### Kernel fuzzing practicalities784785- Use `KCFI` and `KASAN` builds for safer fuzzing; `CONFIG_DEBUG_INFO_BTF=y` helps symbolization.786- Prefer TUN/TAP packet injection over raw sockets for stable replay.787- Stabilize coverage with `kcov` filters and `syz_cover_filter`.788789- **Performance Considerations**:790 - For ARM64 Macs, thermal throttling can significantly reduce execution rates791 - Monitor execution rates for performance degradation792 - VM acceleration and hardware features can improve performance793 - Consider using Hardware Tag-Based KASAN for better performance vs. coverage tradeoff794 - **Linux 6.8+ configs**795 - Enable: `CONFIG_KASAN=y` (or HWASAN on AArch64), `CONFIG_KCSAN=y`, `CONFIG_UBSAN=y`, `CONFIG_KCFI=y`, `CONFIG_DEBUG_INFO_BTF=y`.796 - Prefer `CONFIG_KFENCE=n` during fuzzing; enable later to confirm bugs.797798### AFL799800- Crafting high quality harness801 - identify existing harnesses [oss-fuzz](https://github.com/google/oss-fuzz)802 - adapt and enhance for your fuzzing objectives803- Corpus804 - use `afl-tmin` and `afl-cmin` to tune corpus dynamically805- Code Coverage806 - use `afl-cov` to understand the code coverage807 - then analyze and fine tune your harness808 - and optimize your test corpus based on coverage feedback809- Efficient Crash Triage810 - use [AFLTriage](https://github.com/quic/AFLTriage) and [AddressSanitizer](https://github.com/google/sanitizers/wiki/addresssanitizer)811 - try `afl-collect` / `afl-plot` and enable `ASAN_OPTIONS=abort_on_error=1:symbolize=1`812813#### Modern AFL++ tips814815- Use `-M`/`-S` for parallel fuzzer instances; combine `-c 0` (cmplog) on some slaves.816- Persistent mode (`__AFL_LOOP(N)`) for hot loops; in‑process fuzzing with `afl++-unicorn` for emulated targets.817- Dictionaries (`-x`) and `laf-intel`/`split-switches` to simplify hard conditions at compile time.818819### MacOS820821#### IPC Fuzzing822823- We can mutate and fuzz the message that `mach_msg` is sending824 - Simple to do but slow825 - Hard to determine which process caused the crash826 - Hard to identify code coverage827- We can directly send a message to Target message handler(skipping kernel)828 - Very fast and easy to instrument829 - Easy to understand what caused the crash830 - but Different from end exploit, might need to invoke initialization routines831- Write a Fuzzing harness832833```c834void *lib_handle = dlopen("libexample.dylib", RTLD_LAZY);835pFunction = dlsym(lib_handle, "DesiredFunction");836```837838### EDR839840#### EDR Attack Surface Overview841842EDR solutions present a significant attack surface due to their complex architecture:843844**Potential vulnerability categories:**845846- Memory corruptions in file scanning & emulation (e.g., CVE-2021-1647)847- Arbitrary file deletion via symlink vulnerabilities848- User-Mode IPC authorization issues or memory corruptions849- User-to-driver authorization bypass850- Classic driver vulnerabilities (WDM, KMDF, Mini-Filter)851- Server-to-agent authentication flaws852- Parsing bugs in server responses853- Classic Windows application vulnerabilities (DLL hijacking, file permissions)854- Emulation/sandbox escapes855- Logic bugs in security implementations856857#### Microsoft Defender's Scanning Engine (mpengine.dll)858859Microsoft Defender includes a local analysis engine `mpengine.dll` that performs static checks and uses emulation environments for different file types. This engine presents a significant attack surface due to its complexity and the variety of file formats it processes.860861**Target Characteristics:**862863- Installed by default on Windows systems864- Runs as SYSTEM with high privileges865- Has 1-click remote attack surface through file scanning866- Processes numerous file formats with complex parsing logic867- Prone to memory corruption vulnerabilities868869**Fuzzing Methodology:**870871_Snapshot Fuzzing with WTF:_872873- Uses Windows Terminal Framework for coverage-guided fuzzing874- Takes snapshots after Defender initiates file scanning875- Maintains identical configuration to real environment876- Avoids limitations of manual engine bootstrapping877878_Alternative Approaches:_879880- **kAFL/NYX**: Additional fuzzing frameworks tested881- **Jackalope**: Cross-platform coverage-guided fuzzer882- **Manual harness**: Historic approach using `RSIG_BOOTENGINE` and `RSIG_SCAN_STREAMBUFFER`883884**Practical Attack Scenarios:**885886_Web-based DoS:_887888```html889<!-- Crash Defender via malicious file download -->890<a href="crash.pdf">Download PDF</a>891<!-- Crashes MsMpEng.exe when file is scanned -->892```893894_Network Share DoS:_895896```powershell897# Upload crash file to SMB share before credential dumping898copy crash.doc \\target\share\899# Defender crashes when scanning uploaded file900mimikatz.exe privilege::debug sekurlsa::logonpasswords901```902903**Fuzzing Setup Requirements:**904905_Snapshot Creation:_906907```cpp908// Example WTF harness setup909if (!g_Backend->SetBreakpoint("nt!KeBugCheck2", [](Backend_t *Backend) {910 const uint64_t BCode = Backend->GetArg(0);911 const std::string Filename = fmt::format("crash-{:#x}", BCode);912 Backend->Stop(Crash_t(Filename));913}))914```915916_Coverage Analysis:_917918- Use IDA Lighthouse for visualization919- Monitor for DRIVER_VERIFIER_DETECTED_VIOLATION (0xc4)920- Track IRQL_NOT_LESS_OR_EQUAL (0xa) crashes921922**Defense Implications:**923924- Demonstrates need for robust input validation925- Shows importance of crash recovery mechanisms926- Highlights risks of complex file parsing engines927- Suggests value of sandboxing scanning processes928929##### Cross-platform mpengine.dll Fuzzing930931Recent advances enable fuzzing the latest Windows Defender engine (v1.1.25020.1007) on Linux using **loadlibrary** with Intel PT coverage:932933```cpp934// Disable Lua VM to avoid stability issues935void my_lua_exec(){936 return;937}938939int main(int argc, char** argv){940 // Hook luaV_execute to bypass Lua signature processing941 insert_function_redirect((void*)luaV_execute_address, my_lua_exec, HOOK_REPLACE_FUNCTION);942943 // Setup persistent fuzzing loop944 for (;;) {945 size_t len;946 uint8_t *buf;947 HF_ITER(&buf, &len);948949 ScanDescriptor.UserPtr = fmemopen(buf, len, "r");950951 if (__rsignal(&KernelHandle, RSIG_SCAN_STREAMBUFFER, &ScanParams, sizeof ScanParams) != 0) {952 // Handle scan results953 }954 }955}956```957958##### Performance Optimizations959960- **Persistent Mode**: Eliminates initialization overhead, achieving hundreds of exec/s961- **Intel PT Coverage**: Hardware-based tracing for binary-only targets962- **Lua VM Bypassing**: Reduces crashes and focuses on native code vulnerabilities963964##### honggfuzz with Intel PT Setup965966```bash967# Non-persistent mode (slower, ~4s/execution)968../honggfuzz/honggfuzz -i ~/input/ -W ~/workspace/ --linux_perf_ipt_block -t 10 -- ./mpclient_x64 ___FILE___969970# Per971972…(truncated)