Mantis Pipeline Designer (/mantis-pipeline-adapter)
System Goal
Interactive Pipeline Design Consultant. Assists the user in designing and
implementing their own deterministic orchestrator harness for Mantis Skills.
Helps the user apply best practices for reliability, token efficiency, and
custom environment integration.
Command Definition
- Command:
/mantis-pipeline-adapter
- Description: Interactively guides the design and implementation of custom
deterministic orchestrator harnesses.
Input/Output Contract
- Reads:
workspace/.mantis_state.json (to track current loop pass).
workspace/.mantis_state.json fields active_snapshot, snapshot_history,
and vcs_info.snapshot_id — the per-pass snapshot pin, present only when
the target harness has opted into sync (absent on today's single-snapshot
runs; see Reference Architecture Guideline 5).
schema.json (as the canonical pipeline specification reference).
workspace/findings/*.json (as the State Store).
workspace/learnings.jsonl (to understand memory rotation).
- User's interactive configuration input.
- Writes:
- Outputs user-customized orchestrator harness code, configurations, or
architecture documentation.
- Preconditions:
- User initiates interactive design session.
- Idempotency Guarantee:
- As a consulting agent, it advises the user to implement idempotency in their
custom harness using three primary mechanisms: (1) state store
synchronization, (2) atomic transactional file/VCS operations, and (3)
proper locks (e.g. database/file level locks).
Instructions
Interactively guide the user in designing and building a deterministic pipeline
that wraps Mantis Skills.
Follow these guidelines during the consultation:
- Understand User Context: Ask about their target programming language,
agent framework (if any), execution environments (VMs, local containers,
physical hardware), and scale requirements.
- Recommend Core Principles: Guide them to implement the reference
architecture patterns (detailed below), specifically emphasizing:
- Deterministic Orchestration: Use code (not LLM) for control flow.
- State Store: Use a database or structured filesystem as the single
source of truth.
- Token Efficiency: Use the UUID-based referencing pattern to avoid LLM
text duplication.
- Custom Environment Integration: Use Custom MCP servers for isolated
testing (VMs) or hardware interaction.
- Ensure Schema Consistency: Advise the user to strictly adhere to the
inter-stage data contracts defined in schema.json when
building their harness.
- Adaptive Design: Help them draft the code/architecture tailored to their
specific stack, rather than imposing a rigid template.
- Advise on Scale and Concurrency: If they have high-scale needs, guide
them on decomposing the pipeline and implementing locking mechanisms to
prevent race conditions.
- Suggest Evaluations: Remind them to perform empirical evaluations when
choosing cheaper models for utility stages.
- Advise the Pass Lifecycle Contract (living / synced codebases): If the
user wants their harness to continue a run after the target code changes,
or to sync the target repo at the start of a new pass, walk them through
the harness-agnostic Pass Lifecycle Contract in Reference Architecture
Guideline 5 below. Emphasize that this support is opt-in: a harness that
does not implement the contract MUST leave
snapshot_pinned unset, which
preserves today's single-snapshot behavior byte-for-byte. When --sync is
requested, the harness PINs in the PIN step and passes
--snapshot_root/--snapshot_id normally; Block A (Locator Resolution) is
universal across all code-reading stages.
- Advise on Semantic Retrieval at Scale: If the user is targeting a large
codebase (e.g., thousands of source files, multi-pass campaigns, or multiple
teams contributing findings), walk them through the optional semantic
retrieval patterns in Reference Architecture Guidelines 6 and 7 below.
Emphasize that these are opt-in: they augment the pipeline via a
dedicated query skill or MCP tools, but never modify the existing skills'
own deterministic logic or fail-safe invariants.
- Advise on SAST Seeding: If the user wants to augment LLM-based discovery
with external SAST tool findings (CodeQL, Semgrep, etc.), walk them through
the optional SAST seeding pattern in Reference Architecture Guideline 8
below. Emphasize that this is opt-in: it ingests external findings as
candidates that must earn their verdict through unchanged downstream gates,
and it follows exactly the RAG pattern (provenance-tracked, snapshot-aware,
fallback on failure).
- Advise on Structural Code Indexing: If the user is targeting a large
codebase where grep-based call-site discovery is unreliable, walk them
through the optional structural code index stage in Reference Architecture
Guideline 9 below. Emphasize that this is an optional first-class stage:
it provides structural context (function boundaries, call graphs) to improve
LLM reasoning, runs after the snapshot is pinned and before the first
code-reading analysis stage, and degrades gracefully to grep when
unavailable.
- Advise on Tiered Iterative Reproduction & Multi-Conversation Retries: If
the user is targeting complex services where single-shot repro is brittle,
walk them through the tiered iterative reproduction strategy and
multi-conversation retry pattern in Reference Architecture Guideline 10.
Reference Architecture Guidelines
Use the following guidelines as your technical reference when advising the user.
Core Principles
- Deterministic Orchestration: Do not let the LLM decide the control flow
of the pipeline. Use a programmatic harness to call skills sequentially or in
parallel.
- State on Disk / Database: Use the filesystem
(
workspace/findings/*.json) or a database as the single source of truth.
Skills should read from and write to this store. For horizontal scaling,
recommend a centralized database.
- Deterministic Reporting: Treat findings as internal state. Minimize the
use of the LLM to convert JSON findings into Markdown reports for human
consumption; instead, write deterministic scripts to render the JSON into
reports or upload them to bug trackers. Only use an LLM for non-deterministic
subsets of this (like textual synthesis), such as by providing an executive
summary if necessary.
- Token Efficiency & Reusable Deterministic Tools: Structure LLM outputs to
return only the minimum necessary information (e.g., UUIDs, status codes).
Do not force the LLM to write one-off scripts (e.g., Python or bash) on the
fly for routine tasks like appending JSON fields or merging findings, as this
wastes reasoning tokens. Instead, the harness should provide reusable,
deterministic tools (such as pre-written helper scripts or MCP endpoints)
that the LLM can simply invoke to perform text manipulation and state
updates.
- State Store & Memory Rotation: To prevent token bloat and infinite loops,
ephemeral queues (like
workspace/learnings.jsonl) must be rotated. Upon
successful completion and verification of the Knowledge Base synthesis stage,
the orchestrator should ensure the archive directory exists (e.g.,
mkdir -p workspace/archive/learnings/) and move
workspace/learnings.jsonl to a numbered archive (e.g.,
workspace/archive/learnings/learnings_pass_${N}_${X}.jsonl where ${N} is
the loop pass and ${X} is a sub-index). If the synthesis fails, the active
queue must be left intact to prevent data loss.
Architectural Overview
graph TD
Harness[Programmatic Harness / Orchestrator] <--> DB[(State Store: Disk/DB)]
subgraph Stages [Decomposed Stages]
KB[KB Architect]
TM[Threat Modeler]
P[Plan]
R[Researcher]
D[Deduplicator]
V[Validator/Review]
C[Critic]
Rep[Reproducer]
Ch[Chainer]
Pat[Patcher]
Cal[Calibrator]
Ref[Reflector]
end
Harness --> KB
Harness --> TM
Harness --> P
Harness --> R
Harness --> D
Harness --> V
Harness --> C
Harness --> Rep
Harness --> Ch
Harness --> Pat
Pat -.->|Re-attack Bypass Loop| Rep
Harness --> Cal
Harness --> Ref
subgraph LLM Pool [Tailored LLMs]
ModelA[Frontier Model: Deep Reasoning]
ModelB[Flash/Lite Model: Fast & Cheap]
ModelC[Alternative Provider: Diversified Logic]
end
KB -.-> ModelA
TM -.-> ModelB
P -.-> ModelB
R -.-> ModelA
R -.-> ModelC
D -.-> ModelB
V -.-> ModelB
C -.-> ModelA
Rep -.-> ModelA
Ch -.-> ModelA
Pat -.-> ModelA
Cal -.-> ModelB
Ref -.-> ModelB
1. UUID-Based Referencing Pattern
To prevent the LLM from repeating large blocks of text (which increases latency,
cost, and the risk of mangling data), use UUIDs as the primary key for all
findings.
A. Researcher Stage
- Action: Sweeps the codebase and identifies potential vulnerabilities.
- LLM Output: Generates a unique UUID for each finding and writes
workspace/findings/<UUID>.json containing the full details (matching the
standard schema in Mantis Researcher).
B. Deduplication Stage (Optimized)
Instead of asking the LLM to read all findings, merge them in context, and write
them back, use the following pattern:
Harness Action: Reads all workspace/findings/*.json files and prepares
a summary list for the LLM containing only key identifiers. To align with the
standard schema, map the code_paths array (which uses "file:line" format)
to a simplified summary for the LLM:
[ { "id": "UUID", "file": "path", "line": 12, "snippet": "..." } ].
LLM Action: Analyzes the summary and outputs a mapping of duplicates:
{
"primary_uuid_1": ["duplicate_uuid_a", "duplicate_uuid_b"],
"primary_uuid_2": []
}
Harness Action (Deterministic):
- Reads the content of the affected files.
- Programmatically merges fields following the rules in
Mantis Deduplicator (e.g., union of
code_paths, taking highest severity, concatenating history).
- Updates
workspace/findings/primary_uuid_1.json on disk.
- Ensures the trash directory exists (e.g.,
mkdir -p workspace/findings/.trash/).
- Moves
workspace/findings/duplicate_uuid_a.json and
workspace/findings/duplicate_uuid_b.json to the trash staging directory
(workspace/findings/.trash/).
C. Validation & Review Stages (Reviewer, Critic)
- Harness Action: For each finding
workspace/findings/<UUID>.json, pass
only the relevant code context and finding description to the LLM.
- LLM Action: Output only a structured verification result (e.g.,
{"valid": true, "reason": "..."}).
- Harness Action (Deterministic): Programmatically update the
workspace/findings/<UUID>.json file with the validation status and reason.
2. Adaptable Reproducers via Custom MCP
When validating findings, the agent may need to interact with diverse
environments (VMs, physical hardware). Use the Model Context Protocol (MCP)
to expose a clean, restricted API.
- Architecture:
[Reproducer Agent] <--- MCP ---> [Custom MCP Server] <--- API ---> [Target Env]
- Custom Environments:
- VMs: Implement tools like
reboot_vm(), execute_payload().
- Hardware/USB: Implement tools like
power_cycle_device() (via smart
plug), send_usb_packet().
- Integration Note: If the user's harness uses raw LLM APIs (e.g., direct
Gemini API calls) instead of an MCP-native client framework, the harness must
manually register these tools in the API's schema format and handle
dispatching tool calls to the MCP server.
3. Decomposition & Multi-Model Strategy
A. Pipeline Decomposition & Concurrency
The pipeline can be split into independent services. When scaling horizontally
(e.g., multiple workers running the Reproducer stage in parallel):
- Concurrency Control: Implement database or file locking to ensure two
workers do not attempt to process or update the same finding simultaneously.
- Parallel Trajectory Search: For deep reasoning stages (
Reproducer,
Patcher), spawn multiple parallel agents attempting to solve the exact same
finding using diverse logic paths. For the Reproducer stage, prune all other
trajectories as soon as one worker succeeds to save compute costs while
escaping LLM "give up" loops. For the Patcher stage, wait for all patches to
be generated and tested, then evaluate the successful ones to select the most
minimal, idiomatic, and correct fix.
B. Heterogeneous LLM Selection (Multi-Model)
Match task complexity with the appropriate model tier:
- Frontier Models: For deep reasoning (Research, Reproduce, Patch).
- Flash/Lite Models: For structured utility tasks (Dedupe, Calibrate).
- Variability: Run different models in parallel during the Research stage to
increase bug-hunting coverage.
C. Importance of Evaluation
Emphasize that using cheaper models for utility stages (like deduplication or
calibration) must be validated with empirical evaluations against a benchmark
dataset to ensure quality is not degraded.
4. The Planning Stage and workspace/plan.json
The planning stage plays a critical role in structuring the security campaign.
The strategist (/mantis-plan) generates workspace/plan.json to define
targeted investigations, context pointers, and specific questions for the
auditor. The researcher (/mantis-researcher) reads workspace/plan.json at
startup to guide its sweep. By decoupling strategy and execution via this
structured contract, the orchestrator can easily direct subagents, parallelize
sweeps, and maintain historical context across pipeline runs without repeating
work.
5. The Pass Lifecycle Contract (Living / Synced Codebases)
A custom orchestrator (a bespoke CLI, an ADK agent, an MCP-native pipeline, or
any deterministic harness) does not inherit the living-project lifecycle
that mantis-meta-agent implements. To support continue-after-edits and
opt-in boundary sync without producing silent wrong results (false
VERIFIED_SECURE, false failed_to_reproduce, dropped regressions), the
harness must implement the following harness-agnostic contract. This is the same
contract recorded in schema.json under Non-JSON Contracts;
the Block A–Block G and SNAPSHOT_ID references below name mechanisms each
Mantis stage already carries in its own SKILL.md.
Mantis runs under multiple harnesses (various CLIs, ADK, custom deterministic
pipelines), so the lifecycle must not live only in mantis-meta-agent. Any
harness is conformant iff, per pass, it:
- SYNCs first (Block C) — the very first action; never mid-pass.
- Detects
vcs_info + computes SNAPSHOT_ID (Block D steps 1-5) — only
after sync.
- PINs the immutable copy + writes the sentinel + appends
snapshot_history (Block D step 5, not RECORD).
- Records
vcs_info (incl. snapshot_id) + active_snapshot. Never
record an id or pin before syncing.
- Runs every stage with
--snapshot_root=<SNAPSHOT_ROOT> --snapshot_id=<SNAPSHOT_ID> --state_root=<workspace parent>.
- Archives & increments (existing Stage 15); retried findings keep their
original
discovery_commit.
A harness that does not implement the contract MUST leave snapshot_pinned
unset → today's behavior. When --sync is requested, the harness PINs in the
PIN step and passes --snapshot_root/ --snapshot_id normally; Block A
(Locator Resolution) is universal across all code-reading stages.
Advisory notes when helping a builder implement this contract
- Opt-in, default off. Sync/pinning is a feature the builder turns on. A
harness that never sets
snapshot_pinned behaves exactly like today (one live
snapshot per run). Downstream stages treat an absent
active_snapshot/discovery_commit as the conservative branch, so an
un-upgraded harness is always safe — just not living-project-aware. Do not
advise treating these absent fields as an error.
- Store snapshots OUTSIDE
workspace/. The pinned copy (SNAPSHOT_ROOT)
must live under <state_root>/.mantis_snapshots/pass_<N> (or a clean-VCS
worktree/archive), and its path must not contain the segment /workspace/
— otherwise mantis-patch's state-vs-code path guard misfires. Keep the last
2 snapshots and garbage-collect older ones with the matching teardown
(rm -rf for copies, git worktree remove/prune for worktrees).
- Non-destructive sync only. Sync is the first action of a pass,
never mid-pass, and must be skipped when the tree is dirty, ahead of
upstream, detached, or has no upstream. The harness must never run
git reset --hard, git checkout -- ., git clean, or hg update -C, or
any command that discards uncommitted/untracked/local-commit state — user
edits and in-progress work must survive every pass.
- Full-fidelity
SNAPSHOT_IDs, including dirty / no-VCS. Compute the id
over the whole pinned copy: clean git/hg → commit_hash; dirty git/hg →
commit_hash + ":" + content_hash; multi-vcs →
revision + ":" + content_hash; no-VCS / unknown copyable tree →
"content:" + content_hash. The embedded content hash is exactly what lets an
unchanged dirty or no-VCS tree MATCH across passes and still receive
verification + dedup — and what makes a repo sync that advances commits
under an unchanged manifest revision compare unequal. Never trust a bare
branch name or manifest revision string as an identity.
- Pass the three roots to EVERY stage. Include the findings-only stages
(report, calibrate, reflect): they do not read target code, but they still
read
active_snapshot for provenance/annotation. When the harness archives
and increments, retried findings must keep their original
discovery_commit.
Conformance scenarios
The scenarios below expose nearly every issue in the snapshot model. They are
reference checks, not features: the harness is responsible for preventing or
handling each one in its own environment. The table is a quick-reference; prose
detail follows for each scenario. The State column uses the 3-STATE RULE
(MODE-OFF / HALT / PINNED, branched on active_snapshot presence — see the
global backward-compat rule in schema.json and the advisory
notes above); SNAPSHOT_ID formats follow the ladder in the advisory notes
above (e.g. live:<ts> signals an unpinned/HALT pass).
Invariant legend (the labels below name safety properties enforced by the
blocks and the global backward-compat rule in schema.json):
| Label |
Property |
Enforced by |
| INV-1 |
No false VERIFIED_SECURE |
Block G + HALT ceiling |
| INV-2 |
No false failed_to_reproduce |
Block F + HALT ceiling |
| INV-3 |
No dropped regression |
Block B NOT_MATCHED + POSSIBLE REGRESSION |
| INV-4 |
Within-pass consistency |
Block A sentinel + single pinned snapshot |
| INV-5 |
No user data loss |
Block C non-destructive sync + Block A step 4 |
| INV-6 |
Fail-safe on missing data |
Global backward-compat rule |
Quick-reference table:
| # |
Scenario |
State |
Harness behavior |
Stage behavior |
Block / INV |
Key fields |
| 1 |
Colocated state |
PINNED |
HALT-and-yield (safe default), or relocate state_root outside CODE_ROOT when explicitly authorized (e.g. --auto_relocate_state); SNAPSHOT_ROOT path must not contain /workspace/ |
mantis-patch state-vs-code guard misfires; Block A step 3 confuses SNAPSHOT- vs STATE-relative paths |
A:3, D:3; INV-5 |
active_snapshot.root, snapshot_root, state_root |
| 2 |
Stale active_snapshot (active_snapshot.pass != state.pass_number) |
PINNED → STOP or HALT-degrade |
Block D step 0: handles same-pass re-entry only; if dir missing → STOP, yield to user |
Block A step 2 sentinel may still MATCH (dir retained); CURRENT-PASS CHECK (active_snapshot.pass == state.pass_number) required: mismatch → STOP or HALT-degrade (Block B NOT_MATCHED, no authoritative verdicts) |
A:2, D:0, B; INV-1, INV-3, INV-4, INV-6 |
active_snapshot.{root, snapshot_id, snapshot_pinned, pass}, state.pass_number, discovery_commit |
| 3 |
Pin failure |
HALT |
Block D step 2/4: skip copy on ENOSPC/error → step 5b; still write active_snapshot + pass roots |
Authoritative verdicts forbidden; Block B always NOT_MATCHED; reproduce not_attempted; patch VERIFICATION_INCOMPLETE |
D:2, D:4, D:5b; INV-1, INV-2, INV-6 |
active_snapshot.{snapshot_id, snapshot_pinned} |
| 4 |
Patched shadows |
PINNED (pass); --snapshot_pinned=false arg |
Pass --target_root=<PATCHED_SHADOW_ROOT> + --snapshot_pinned=false to reattack sub-agent |
Block A step 1a: CODE_ROOT=--target_root (authoritative); step 2 sentinel SKIPPED (sentinel-EXEMPT) |
A:1a, A:2; INV-4 |
target_root, snapshot_pinned (arg), snapshot_root, discovery_commit |
| 5 |
Different-snapshot duplicate candidates |
PINNED |
No special action — both passes pinned correctly; dedupe handles it |
Block B pairwise: discovery_commit differs → NOT_MATCHED → keep ACTIVE + possible_duplicate_of; POSSIBLE REGRESSION if archived was RESOLVED |
B; INV-3, INV-6 |
discovery_commit, possible_duplicate_of, status, patch_status |
| 6 |
Absent sink evidence |
Any |
No special action — Block F is a stage-level mechanical gate |
Block F: evidence absent (build error, exit 127, sink unreached) → not_attempted (retry-eligible), NEVER failed_to_reproduce; HALT ceiling additionally forces not_attempted |
F; INV-2, INV-6 |
repro_status, reattack_status, repro_hints |
Per-scenario detail:
1. Colocated state (state_root nested inside CODE_ROOT / snapshot root)
— The pinned SNAPSHOT_ROOT must live under
<state_root>/.mantis_snapshots/pass_<N> (or a clean-VCS worktree/archive), and
its path must not contain the segment /workspace/ — otherwise
mantis-patch's state-vs-code path guard misfires (state files appear to be
"under CODE_ROOT"). If state_root itself is inside CODE_ROOT, the harness
must HALT-and-yield (safe default) or, when explicitly authorized (e.g.
--auto_relocate_state), relocate it outside the snapshot before pinning. Block
A step 3 distinguishes SNAPSHOT-RELATIVE path fields (read under CODE_ROOT)
from STATE-RELATIVE fields (read under state_root/workspace, never prefixed
with CODE_ROOT); colocation breaks this separation.
2. Stale active_snapshot (active_snapshot.pass != state.pass_number —
active_snapshot was preserved across the Stage 15 pass increment) — Block D
step 0 (crash-resume) handles only the SAME-pass re-entry case
(active_snapshot.pass == N → reuse). It does NOT catch a stale snapshot
carried across the Stage 15 pass increment, because Stage 15 deliberately
preserves active_snapshot while bumping pass_number (see Stage 15). Two
sub-cases:
(a) The prior snapshot dir is now MISSING: Block D step 0 STOPs and yields to
the user (never re-pin to a possibly-drifted live tree). (b) The prior snapshot
dir still EXISTS (default keep-2 retention) and its sentinel matches the
preserved active_snapshot.snapshot_id: Block A step 2 sentinel check SUCCEEDS
(it only compares the sentinel file to SNAPSHOT_ID, not to the current pass).
Block B's pairwise discovery_commit check would MATCH a carried-forward
finding against a new finding stamped with the same stale SNAPSHOT_ID,
silently dropping it as DUPLICATE — a false authoritative verdict.
To prevent (b), the HARNESS MUST guarantee that
active_snapshot.pass == state.pass_number before any consumer stage reads it.
The reference harness (mantis-meta-agent) satisfies this by re-pinning every
pass (Block D step 0 sees active_snapshot.pass != N → re-pins → refreshes
active_snapshot.pass before any stage runs), so sub-case (b) never fires
there. A custom harness that preserves active_snapshot across the Stage 15
pass increment WITHOUT re-pinning MUST either (a) re-pin every pass (the
reference behavior), or (b) inject an equivalent pre-stage gate that refreshes
active_snapshot.pass or clears active_snapshot entirely before invoking
stages. Stages CANNOT self-detect this staleness via Block B (which is
snapshot_id-only, not pass-aware): a carried-forward finding and a new
finding stamped with the same stale SNAPSHOT_ID will MATCH in Block B despite
the snapshot being stale. The active_snapshot.pass field is defined in
schema.json #/$defs/state/active_snapshot/pass for exactly this check. The
harness's Block D step 0 reuse check is NOT a substitute: it only fires on
same-pass re-entry. (Stages that read active_snapshot MAY additionally
self-check defensively — see each stage's Step 0 sentinel check — but the
binding guarantee is on the harness.)
3. Pin failure (snapshot copy fails — disk full, permissions, too-large
tree) — Block D step 2 (free-space precheck): compare du -s of the live tree
to df free space at state_root; if it won't fit → skip copy → step 5b. Block
D step 4 (failure-tolerant verify): check copy exit status + sanity check (file
count/size within ~90%); on failure → step 5b (unpinned/HALT). Step 5b:
SNAPSHOT_ROOT=<live root>, snapshot_pinned=false,
SNAPSHOT_ID="live:"+ISO8601. The harness still writes active_snapshot and
still passes --snapshot_root/--snapshot_id to stages so they see the HALT
signal. Every stage then degrades conservatively: authoritative verdicts
forbidden (VERIFIED_SECURE, failed_to_reproduce, DUPLICATE,
FALSE_POSITIVE, NON_VIABLE); Block B always returns NOT_MATCHED; reproduce
records not_attempted; patch's best attainable is VERIFICATION_INCOMPLETE.
4. Patched shadows (--target_root pointing at a pre-mutated tree;
sentinel-exempt path 1a in Block A) — mantis-patch passes
--target_root=<PATCHED_SHADOW_ROOT> and --snapshot_pinned=false to the
reproduce sub-agent for re-attack verification. Block A step 1a:
CODE_ROOT = --target_root (authoritative override, overrides --snapshot_root
and state fallback). Block A step 2: sentinel check SKIPPED (a --target_root
tree is deliberately mutated and is sentinel-EXEMPT). The
--snapshot_pinned=false argument is the sentinel-exemption, NOT a HALT signal
— detect HALT by reading STATE (active_snapshot.snapshot_id starts with
live:, equivalently active_snapshot.snapshot_pinned is false in state),
never from the argument passed on this invocation. The finding's
discovery_commit is unaffected — it retains the pass-level SNAPSHOT_ID from
when it was discovered; only the --snapshot_pinned=false argument is local to
the reattack invocation.
5. Different-snapshot duplicate candidates (cross-pass dedupe where
discovery_commit differs — the pairwise Block B NOT_MATCHED path) — Both
passes pinned correctly; the findings simply come from different snapshots.
mantis-dedupe Block B pairwise check compares the CURRENT finding's
discovery_commit against the ARCHIVED finding's discovery_commit (NOT
against the global SNAPSHOT_ID). If they differ → NOT_MATCHED. NOT_MATCHED
keeps the current finding ACTIVE and sets possible_duplicate_of (a soft,
non-terminal hint — the finding is NOT filtered or trashed). If the archived
finding was RESOLVED (patch_status in {VERIFIED_SECURE,
MITIGATION_PROPOSED} OR status==FALSE_POSITIVE OR
production_viability==NON_VIABLE) AND the pair is NOT_MATCHED → POSSIBLE
REGRESSION: keep ACTIVE, add a history note, never filter (a reverted fix
re-discovered on new code must never be trashed).
6. Absent sink evidence (Block F — PoC compiles but produces no reached-sink
evidence; not_attempted vs failed_to_reproduce) — mantis-reproduce Block
F: if EVIDENCE is ABSENT (any compiler/build nonzero exit, exit 127
command-not-found, exit 2 "No such file", or the sink was never reached) →
repro_status = not_attempted (retry-eligible), STOP. NEVER
failed_to_reproduce. In --reattack mode: leave reattack_status UNSET with
a history note "setup_failed" — NEVER failed_to_bypass. failed_to_reproduce
is reserved for when the harness PROVABLY reached the vulnerable entrypoint —
i.e. reached-sink evidence, not setup evidence — but the bug did not fire.
Reached-sink evidence must originate INSIDE the invoked path or from
target-produced tracing/backtraces: (a) a PoC script/source harness writes
MANTIS_REACHED_ENTRYPOINT to a sidecar file at the point just before the sink
call, within its own execution flow (the marker write is part of the invoked
path, not a pre-launch step); OR (b) for binary/firmware/raw-payload targets,
the captured crash backtrace or sanitizer trace (ASan/UBSan/MSan/TSan)
explicitly names the target sink function (target-produced tracing). A marker
written by an external wrapper BEFORE invoking the target is SETUP EVIDENCE ONLY
(proves "launch attempted," not "sink reached") and does NOT by itself justify
failed_to_reproduce — treat it as EVIDENCE ABSENT for the decision gate.
Evidence is recorded in repro_hints. In HALT mode, the HALT ceiling
additionally forces not_attempted (no failed_to_reproduce), since a negative
result on an unpinned tree cannot be trusted as authoritative.
6. Semantic Retrieval (RAG) for Large Codebases
For small repositories, the planner can manually scan workspace/kb/index.md
and the researcher can grep for call-sites. At scale (thousands of files, deep
directory trees, multi-pass campaigns), these approaches miss relevant context
and waste tokens reading irrelevant files. A semantic retrieval layer lets the
planner and researcher query for relevant KB entries and code locations without
reading everything.
Two implementations are supported, sharing the same data contract:
- Option A (Default — Skill-Based): A dedicated skill that runs a
BM25/TF-IDF helper script over
chunks.jsonl. Zero external dependencies —
works air-gapped, no vector embeddings or vector store required. Optional
vector embedding support if available.
- Option B (Maximum Scale — MCP-Based): The harness owns a persistent vector
index using vector embeddings, serving persistent
semantic_search_kb /
semantic_search_code MCP tools. Better for very large codebases where
per-invocation BM25 is too slow.
Both are opt-in. The existing skills are not modified; the planner and
researcher receive runtime instructions to use whichever retrieval mechanism is
available, falling back to today's manual behavior if neither is present.
Retrieval results are coverage HINTs only — they decide ordering and
prioritization, never the membership of the audit set. A miss must never cause a
file, call-site, or investigation to be skipped or dropped.
A. Shared Data Contract: chunks.jsonl
After Stage 2 (/mantis-architecture) completes, chunks are extracted into
workspace/kb/chunks.jsonl (one JSON object per line). The harness can do this
post-hoc by reading workspace/kb/*.md, or the architecture skill can be
instructed to write it during synthesis as a text-only side effect. Two chunk
types are produced:
KB chunks from the existing workspace/kb/*.md files:
{"id": "auth_module:0", "source_file": "workspace/kb/entities/auth_module.md", "entity_type": "entity", "chunk_text": "The auth module handles..."}
Code chunks from CODE_ROOT (the pinned snapshot). Each chunk includes
the file path and line range so the researcher can request specific files
from the snapshot:
{"id": "src/parser.c:0", "source_file": "src/parser.c", "start_line": 1, "end_line": 80, "chunk_text": "int parse_input(..."}
The first line of chunks.jsonl is a provenance header recording the
SNAPSHOT_ID the chunks were built against:
{"_provenance": true, "snapshot_id": "abc123", "kb_snapshot_id": "abc123"}
Before serving queries, check snapshot_id in the provenance header against the
current SNAPSHOT_ID; rebuild if they differ. In MODE-OFF (no
active_snapshot), kb_snapshot_id is never stamped — skip the index entirely
and let skills fall back to manual scanning. Never build code chunks from the
live tree — they must reflect the pinned copy the skills are reading.
B. Option A: Skill-Based Retrieval (Default — No Infrastructure)
A dedicated skill reads chunks.jsonl and writes+runs a helper script (e.g.
workspace/helpers/search_chunks.py) that performs BM25/TF-IDF similarity
search. The script is generated by the agent at runtime — no code is shipped
with the skill (same pattern as mantis-dedupe's merge_findings.py). This
requires zero external dependencies — no embedding model, no vector store, no
MCP server. It works in air-gapped and VPC-SC environments.
A complete reference blueprint for this skill is available at
references/mantis-kb-query.md. Builders can
adapt it to their environment. The blueprint includes Block A (Locator
Resolution), chunk provenance checking, the versioned helper script contract
(MANTIS_HELPER_VERSION = 1), and the JSON output schema.
- Invocation: The planner or researcher spawns the skill as a sub-agent with
a query string. The skill writes the helper if not already present, runs it,
and returns top-K matching chunks as JSON.
- Optional embeddings: If vector embeddings are available, the agent can be
instructed to use cosine similarity instead of BM25. This is a runtime
configuration toggle, not a different skill.
- Snapshot safety: The skill reads
active_snapshot from state via Block A
(same as every other skill) and checks chunk provenance before serving.
C. Option B: MCP-Based Retrieval (For Maximum Scale)
For very large codebases where per-invocation BM25 is too slow, the harness can
own a persistent vector index using vector embeddings, serving two MCP tools
(following the same pattern as Guideline 2's Custom MCP for VMs/hardware):
semantic_search_kb(query: string) → [{id, source_file, entity_type, chunk_text, score}]
— Searches KB chunks. Returns relevant entity/vulnerability markdown context.
semantic_search_code(query: string) → [{file, start_line, end_line, snippet, score}]
— Searches code chunks from the pinned snapshot. Returns relevant code
locations.
The harness manages the vector index lifecycle: build from chunks.jsonl (or
directly from CODE_ROOT), rebuild when SNAPSHOT_ID changes, and handle
freshness checks. In HALT mode, serve with a STALE flag or refuse. In
MODE-OFF, skip entirely.
D. Per-Skill Augmentation Guidance
When a retrieval mechanism (skill or MCP) is available, instruct the following
skills to use it. These are runtime instructions passed by the harness or
meta-agent when invoking the skill — the skill files themselves are not
modified:
mantis-architecture: No changes needed. The harness chunks the existing
workspace/kb/*.md files after the architect completes Stage 2. If the
builder prefers, they may instruct the architect to also write
workspace/kb/chunks.jsonl during synthesis (Step 3) as a text-only side
effect — but this is optional, since the harness can extract chunks post-hoc.
mantis-plan: If a retrieval mechanism is available, instruct the planner
to use it to discover kb_references for each investigation instead of only
manually scanning workspace/kb/index.md. For each investigation, query with
the investigation title and target file names, then add the top-K matching KB
entity/vulnerability files to the kb_references array. Manual scanning of
index.md remains the fallback when no mechanism is available.
mantis-researcher: If a retrieval mechanism is available, instruct Wave 1
sub-agents to use it to PRIORITIZE relevant call-sites and cross-module data
flows into sinks (e.g., "where does untrusted input reach memcpy in the
parser module"). Semantic search SUPPLEMENTS grep as a ranking HINT ONLY — it
decides ORDER, never MEMBERSHIP
…(truncated)
1---2name: mantis-pipeline-adapter3description: Interactively guides the design and implementation of custom deterministic orchestrator harnesses. Use when a user wants to build their own pipeline to wrap and run Mantis skills reliably. Don't use for executing the default pipeline directly.4---5
6# Mantis Pipeline Designer (/mantis-pipeline-adapter)
7
8## System Goal
9
10Interactive Pipeline Design Consultant. Assists the user in designing and
11implementing their own deterministic orchestrator harness for Mantis Skills.
12Helps the user apply best practices for reliability, token efficiency, and
13custom environment integration.
14
15## Command Definition
16
17- **Command:** `/mantis-pipeline-adapter`
18- **Description:** Interactively guides the design and implementation of custom
19 deterministic orchestrator harnesses.
20
21## Input/Output Contract
22
23- **Reads**:
24 - `workspace/.mantis_state.json` (to track current loop pass).
25 - `workspace/.mantis_state.json` fields `active_snapshot`, `snapshot_history`,
26 and `vcs_info.snapshot_id` — the per-pass snapshot pin, present only when
27 the target harness has opted into sync (absent on today's single-snapshot
28 runs; see Reference Architecture Guideline 5).
29 - `schema.json` (as the canonical pipeline specification reference).
30 - `workspace/findings/*.json` (as the State Store).
31 - `workspace/learnings.jsonl` (to understand memory rotation).
32 - User's interactive configuration input.
33- **Writes**:
34 - Outputs user-customized orchestrator harness code, configurations, or
35 architecture documentation.
36- **Preconditions**:
37 - User initiates interactive design session.
38- **Idempotency Guarantee**:
39 - As a consulting agent, it advises the user to implement idempotency in their
40 custom harness using three primary mechanisms: (1) state store
41 synchronization, (2) atomic transactional file/VCS operations, and (3)
42 proper locks (e.g. database/file level locks).
43
44## Instructions
45
46Interactively guide the user in designing and building a deterministic pipeline
47that wraps Mantis Skills.
48
49Follow these guidelines during the consultation:
50
5101. **Understand User Context:** Ask about their target programming language,
52 agent framework (if any), execution environments (VMs, local containers,
53 physical hardware), and scale requirements.
5402. **Recommend Core Principles:** Guide them to implement the reference
55 architecture patterns (detailed below), specifically emphasizing:
56 - **Deterministic Orchestration**: Use code (not LLM) for control flow.
57 - **State Store**: Use a database or structured filesystem as the single
58 source of truth.
59 - **Token Efficiency**: Use the UUID-based referencing pattern to avoid LLM
60 text duplication.
61 - **Custom Environment Integration**: Use Custom MCP servers for isolated
62 testing (VMs) or hardware interaction.
6303. **Ensure Schema Consistency**: Advise the user to strictly adhere to the
64 inter-stage data contracts defined in [schema.json](../schema.json) when
65 building their harness.
6604. **Adaptive Design**: Help them draft the code/architecture tailored to their
67 specific stack, rather than imposing a rigid template.
6805. **Advise on Scale and Concurrency**: If they have high-scale needs, guide
69 them on decomposing the pipeline and implementing locking mechanisms to
70 prevent race conditions.
7106. **Suggest Evaluations:** Remind them to perform empirical evaluations when
72 choosing cheaper models for utility stages.
7307. **Advise the Pass Lifecycle Contract (living / synced codebases):** If the
74 user wants their harness to *continue a run after the target code changes*,
75 or to *sync the target repo at the start of a new pass*, walk them through
76 the harness-agnostic Pass Lifecycle Contract in Reference Architecture
77 Guideline 5 below. Emphasize that this support is **opt-in**: a harness that
78 does not implement the contract MUST leave `snapshot_pinned` unset, which
79 preserves today's single-snapshot behavior byte-for-byte. When `--sync` is
80 requested, the harness PINs in the PIN step and passes
81 `--snapshot_root`/`--snapshot_id` normally; Block A (Locator Resolution) is
82 universal across all code-reading stages.
8308. **Advise on Semantic Retrieval at Scale:** If the user is targeting a large
84 codebase (e.g., thousands of source files, multi-pass campaigns, or multiple
85 teams contributing findings), walk them through the optional semantic
86 retrieval patterns in Reference Architecture Guidelines 6 and 7 below.
87 Emphasize that these are **opt-in**: they augment the pipeline via a
88 dedicated query skill or MCP tools, but never modify the existing skills'
89 own deterministic logic or fail-safe invariants.
9009. **Advise on SAST Seeding:** If the user wants to augment LLM-based discovery
91 with external SAST tool findings (CodeQL, Semgrep, etc.), walk them through
92 the optional SAST seeding pattern in Reference Architecture Guideline 8
93 below. Emphasize that this is **opt-in**: it ingests external findings as
94 candidates that must earn their verdict through unchanged downstream gates,
95 and it follows exactly the RAG pattern (provenance-tracked, snapshot-aware,
96 fallback on failure).
9710. **Advise on Structural Code Indexing:** If the user is targeting a large
98 codebase where grep-based call-site discovery is unreliable, walk them
99 through the optional structural code index stage in Reference Architecture
100 Guideline 9 below. Emphasize that this is an **optional first-class stage**:
101 it provides structural context (function boundaries, call graphs) to improve
102 LLM reasoning, runs after the snapshot is pinned and before the first
103 code-reading analysis stage, and degrades gracefully to grep when
104 unavailable.
10511. **Advise on Tiered Iterative Reproduction & Multi-Conversation Retries:** If
106 the user is targeting complex services where single-shot repro is brittle,
107 walk them through the tiered iterative reproduction strategy and
108 multi-conversation retry pattern in Reference Architecture Guideline 10.
109
110## Reference Architecture Guidelines
111
112Use the following guidelines as your technical reference when advising the user.
113
114### Core Principles
115
1161. **Deterministic Orchestration:** Do not let the LLM decide the control flow
117 of the pipeline. Use a programmatic harness to call skills sequentially or in
118 parallel.
1192. **State on Disk / Database:** Use the filesystem
120 (`workspace/findings/*.json`) or a database as the single source of truth.
121 Skills should read from and write to this store. For horizontal scaling,
122 recommend a centralized database.
1233. **Deterministic Reporting:** Treat findings as internal state. Minimize the
124 use of the LLM to convert JSON findings into Markdown reports for human
125 consumption; instead, write deterministic scripts to render the JSON into
126 reports or upload them to bug trackers. Only use an LLM for non-deterministic
127 subsets of this (like textual synthesis), such as by providing an executive
128 summary if necessary.
1294. **Token Efficiency & Reusable Deterministic Tools:** Structure LLM outputs to
130 return only the *minimum necessary information* (e.g., UUIDs, status codes).
131 Do not force the LLM to write one-off scripts (e.g., Python or bash) on the
132 fly for routine tasks like appending JSON fields or merging findings, as this
133 wastes reasoning tokens. Instead, the harness should provide reusable,
134 deterministic tools (such as pre-written helper scripts or MCP endpoints)
135 that the LLM can simply invoke to perform text manipulation and state
136 updates.
1375. **State Store & Memory Rotation:** To prevent token bloat and infinite loops,
138 ephemeral queues (like `workspace/learnings.jsonl`) must be rotated. Upon
139 successful completion and verification of the Knowledge Base synthesis stage,
140 the orchestrator should ensure the archive directory exists (e.g.,
141 `mkdir -p workspace/archive/learnings/`) and **move**
142 `workspace/learnings.jsonl` to a numbered archive (e.g.,
143 `workspace/archive/learnings/learnings_pass_${N}_${X}.jsonl` where `${N}` is
144 the loop pass and `${X}` is a sub-index). If the synthesis fails, the active
145 queue must be left intact to prevent data loss.
146
147### Architectural Overview
148
149```mermaid
150graph TD
151 Harness[Programmatic Harness / Orchestrator] <--> DB[(State Store: Disk/DB)]
152
153 subgraph Stages [Decomposed Stages]
154 KB[KB Architect]
155 TM[Threat Modeler]
156 P[Plan]
157 R[Researcher]
158 D[Deduplicator]
159 V[Validator/Review]
160 C[Critic]
161 Rep[Reproducer]
162 Ch[Chainer]
163 Pat[Patcher]
164 Cal[Calibrator]
165 Ref[Reflector]
166 end
167
168 Harness --> KB
169 Harness --> TM
170 Harness --> P
171 Harness --> R
172 Harness --> D
173 Harness --> V
174 Harness --> C
175 Harness --> Rep
176 Harness --> Ch
177 Harness --> Pat
178 Pat -.->|Re-attack Bypass Loop| Rep
179 Harness --> Cal
180 Harness --> Ref
181
182 subgraph LLM Pool [Tailored LLMs]
183 ModelA[Frontier Model: Deep Reasoning]
184 ModelB[Flash/Lite Model: Fast & Cheap]
185 ModelC[Alternative Provider: Diversified Logic]
186 end
187
188 KB -.-> ModelA
189 TM -.-> ModelB
190 P -.-> ModelB
191 R -.-> ModelA
192 R -.-> ModelC
193 D -.-> ModelB
194 V -.-> ModelB
195 C -.-> ModelA
196 Rep -.-> ModelA
197 Ch -.-> ModelA
198 Pat -.-> ModelA
199 Cal -.-> ModelB
200 Ref -.-> ModelB
201```
202
203### 1. UUID-Based Referencing Pattern
204
205To prevent the LLM from repeating large blocks of text (which increases latency,
206cost, and the risk of mangling data), use UUIDs as the primary key for all
207findings.
208
209#### A. Researcher Stage
210
211- **Action:** Sweeps the codebase and identifies potential vulnerabilities.
212- **LLM Output:** Generates a unique UUID for each finding and writes
213 `workspace/findings/<UUID>.json` containing the full details (matching the
214 standard schema in [Mantis Researcher](../mantis-researcher/SKILL.md)).
215
216#### B. Deduplication Stage (Optimized)
217
218Instead of asking the LLM to read all findings, merge them in context, and write
219them back, use the following pattern:
220
2211. **Harness Action:** Reads all `workspace/findings/*.json` files and prepares
222 a summary list for the LLM containing only key identifiers. To align with the
223 standard schema, map the `code_paths` array (which uses `"file:line"` format)
224 to a simplified summary for the LLM:
225 `[ { "id": "UUID", "file": "path", "line": 12, "snippet": "..." } ]`.
226
2272. **LLM Action:** Analyzes the summary and outputs a mapping of duplicates:
228
229 ```json
230 {
231 "primary_uuid_1": ["duplicate_uuid_a", "duplicate_uuid_b"],
232 "primary_uuid_2": []
233 }
234 ```
235
2363. **Harness Action (Deterministic):**
237
238 - Reads the content of the affected files.
239 - Programmatically merges fields following the rules in
240 [Mantis Deduplicator](../mantis-dedupe/SKILL.md) (e.g., union of
241 `code_paths`, taking highest severity, concatenating history).
242 - Updates `workspace/findings/primary_uuid_1.json` on disk.
243 - Ensures the trash directory exists (e.g.,
244 `mkdir -p workspace/findings/.trash/`).
245 - Moves `workspace/findings/duplicate_uuid_a.json` and
246 `workspace/findings/duplicate_uuid_b.json` to the trash staging directory
247 (`workspace/findings/.trash/`).
248
249#### C. Validation & Review Stages (Reviewer, Critic)
250
251- **Harness Action:** For each finding `workspace/findings/<UUID>.json`, pass
252 only the relevant code context and finding description to the LLM.
253- **LLM Action:** Output *only* a structured verification result (e.g.,
254 `{"valid": true, "reason": "..."}`).
255- **Harness Action (Deterministic):** Programmatically update the
256 `workspace/findings/<UUID>.json` file with the validation status and reason.
257
258### 2. Adaptable Reproducers via Custom MCP
259
260When validating findings, the agent may need to interact with diverse
261environments (VMs, physical hardware). Use the **Model Context Protocol (MCP)**
262to expose a clean, restricted API.
263
264- **Architecture**:
265 `[Reproducer Agent] <--- MCP ---> [Custom MCP Server] <--- API ---> [Target Env]`
266- **Custom Environments**:
267 - *VMs*: Implement tools like `reboot_vm()`, `execute_payload()`.
268 - *Hardware/USB*: Implement tools like `power_cycle_device()` (via smart
269 plug), `send_usb_packet()`.
270- **Integration Note**: If the user's harness uses raw LLM APIs (e.g., direct
271 Gemini API calls) instead of an MCP-native client framework, the harness must
272 manually register these tools in the API's schema format and handle
273 dispatching tool calls to the MCP server.
274
275### 3. Decomposition & Multi-Model Strategy
276
277#### A. Pipeline Decomposition & Concurrency
278
279The pipeline can be split into independent services. When scaling horizontally
280(e.g., multiple workers running the `Reproducer` stage in parallel):
281
282- **Concurrency Control**: Implement database or file locking to ensure two
283 workers do not attempt to process or update the same finding simultaneously.
284- **Parallel Trajectory Search**: For deep reasoning stages (`Reproducer`,
285 `Patcher`), spawn multiple parallel agents attempting to solve the exact same
286 finding using diverse logic paths. For the `Reproducer` stage, prune all other
287 trajectories as soon as one worker succeeds to save compute costs while
288 escaping LLM "give up" loops. For the `Patcher` stage, wait for all patches to
289 be generated and tested, then evaluate the successful ones to select the most
290 minimal, idiomatic, and correct fix.
291
292#### B. Heterogeneous LLM Selection (Multi-Model)
293
294Match task complexity with the appropriate model tier:
295
296- *Frontier Models*: For deep reasoning (Research, Reproduce, Patch).
297- *Flash/Lite Models*: For structured utility tasks (Dedupe, Calibrate).
298- *Variability*: Run different models in parallel during the Research stage to
299 increase bug-hunting coverage.
300
301#### C. Importance of Evaluation
302
303Emphasize that using cheaper models for utility stages (like deduplication or
304calibration) must be validated with empirical evaluations against a benchmark
305dataset to ensure quality is not degraded.
306
307### 4. The Planning Stage and workspace/plan.json
308
309The planning stage plays a critical role in structuring the security campaign.
310The strategist (`/mantis-plan`) generates `workspace/plan.json` to define
311targeted investigations, context pointers, and specific questions for the
312auditor. The researcher (`/mantis-researcher`) reads `workspace/plan.json` at
313startup to guide its sweep. By decoupling strategy and execution via this
314structured contract, the orchestrator can easily direct subagents, parallelize
315sweeps, and maintain historical context across pipeline runs without repeating
316work.
317
318### 5. The Pass Lifecycle Contract (Living / Synced Codebases)
319
320A custom orchestrator (a bespoke CLI, an ADK agent, an MCP-native pipeline, or
321any deterministic harness) does **not** inherit the living-project lifecycle
322that `mantis-meta-agent` implements. To support *continue-after-edits* and
323*opt-in boundary sync* **without producing silent wrong results** (false
324`VERIFIED_SECURE`, false `failed_to_reproduce`, dropped regressions), the
325harness must implement the following harness-agnostic contract. This is the same
326contract recorded in [schema.json](../schema.json) under **Non-JSON Contracts**;
327the `Block A`–`Block G` and `SNAPSHOT_ID` references below name mechanisms each
328Mantis stage already carries in its own `SKILL.md`.
329
330Mantis runs under multiple harnesses (various CLIs, ADK, custom deterministic
331pipelines), so the lifecycle must not live only in `mantis-meta-agent`. Any
332harness is **conformant** iff, per pass, it:
333
3341. **SYNCs first** (Block C) — the very first action; never mid-pass.
3352. **Detects `vcs_info` + computes `SNAPSHOT_ID`** (Block D steps 1-5) — only
336 after sync.
3373. **PINs** the immutable copy + writes the sentinel + **appends**
338 `snapshot_history` (Block D step 5, not RECORD).
3394. **Records** `vcs_info` (incl. `snapshot_id`) + `active_snapshot`. Never
340 record an id or pin before syncing.
3415. **Runs every stage** with
342 `--snapshot_root=<SNAPSHOT_ROOT> --snapshot_id=<SNAPSHOT_ID> --state_root=<workspace parent>`.
3436. **Archives & increments** (existing Stage 15); retried findings keep their
344 **original** `discovery_commit`.
345
346A harness that does not implement the contract MUST leave `snapshot_pinned`
347unset → today's behavior. When `--sync` is requested, the harness PINs in the
348PIN step and passes `--snapshot_root`/ `--snapshot_id` normally; Block A
349(Locator Resolution) is universal across all code-reading stages.
350
351#### Advisory notes when helping a builder implement this contract
352
353- **Opt-in, default off.** Sync/pinning is a feature the builder turns on. A
354 harness that never sets `snapshot_pinned` behaves exactly like today (one live
355 snapshot per run). Downstream stages treat an absent
356 `active_snapshot`/`discovery_commit` as the conservative branch, so an
357 un-upgraded harness is always safe — just not living-project-aware. Do **not**
358 advise treating these absent fields as an error.
359- **Store snapshots OUTSIDE `workspace/`.** The pinned copy (`SNAPSHOT_ROOT`)
360 must live under `<state_root>/.mantis_snapshots/pass_<N>` (or a clean-VCS
361 worktree/archive), and its path **must not** contain the segment `/workspace/`
362 — otherwise `mantis-patch`'s state-vs-code path guard misfires. Keep the last
363 2 snapshots and garbage-collect older ones with the matching teardown
364 (`rm -rf` for copies, `git worktree remove/prune` for worktrees).
365- **Non-destructive sync only.** Sync is the **first** action of a pass,
366 **never** mid-pass, and must be **skipped** when the tree is dirty, ahead of
367 upstream, detached, or has no upstream. The harness must **never** run
368 `git reset --hard`, `git checkout -- .`, `git clean`, or `hg update -C`, or
369 any command that discards uncommitted/untracked/local-commit state — user
370 edits and in-progress work must survive every pass.
371- **Full-fidelity `SNAPSHOT_ID`s, including dirty / no-VCS.** Compute the id
372 over the **whole** pinned copy: clean git/hg → `commit_hash`; dirty git/hg →
373 `commit_hash + ":" + content_hash`; multi-vcs →
374 `revision + ":" + content_hash`; no-VCS / unknown copyable tree →
375 `"content:" + content_hash`. The embedded content hash is exactly what lets an
376 **unchanged dirty or no-VCS tree MATCH across passes** and still receive
377 verification + dedup — and what makes a `repo sync` that advances commits
378 under an unchanged manifest `revision` compare **unequal**. Never trust a bare
379 branch name or manifest revision string as an identity.
380- **Pass the three roots to EVERY stage.** Include the findings-only stages
381 (report, calibrate, reflect): they do not read target code, but they still
382 read `active_snapshot` for provenance/annotation. When the harness archives
383 and increments, retried findings must keep their **original**
384 `discovery_commit`.
385
386#### Conformance scenarios
387
388The scenarios below expose nearly every issue in the snapshot model. They are
389**reference checks**, not features: the harness is responsible for preventing or
390handling each one in its own environment. The table is a quick-reference; prose
391detail follows for each scenario. The **State** column uses the 3-STATE RULE
392(MODE-OFF / HALT / PINNED, branched on `active_snapshot` presence — see the
393global backward-compat rule in [schema.json](../schema.json) and the advisory
394notes above); `SNAPSHOT_ID` formats follow the ladder in the advisory notes
395above (e.g. `live:<ts>` signals an unpinned/HALT pass).
396
397**Invariant legend** (the labels below name safety properties enforced by the
398blocks and the global backward-compat rule in [schema.json](../schema.json)):
399
400| Label | Property | Enforced by |
401| ----- | ------------------------------ | --------------------------------------------- |
402| INV-1 | No false `VERIFIED_SECURE` | Block G + HALT ceiling |
403| INV-2 | No false `failed_to_reproduce` | Block F + HALT ceiling |
404| INV-3 | No dropped regression | Block B NOT_MATCHED + POSSIBLE REGRESSION |
405| INV-4 | Within-pass consistency | Block A sentinel + single pinned snapshot |
406| INV-5 | No user data loss | Block C non-destructive sync + Block A step 4 |
407| INV-6 | Fail-safe on missing data | Global backward-compat rule |
408
409**Quick-reference table:**
410
411| # | Scenario | State | Harness behavior | Stage behavior | Block / INV | Key fields |
412| --- | ------------------------------------------------------------------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------- | ----------------------------------------------------------------------------------------------------- |
413| 1 | Colocated state | PINNED | HALT-and-yield (safe default), or relocate `state_root` outside `CODE_ROOT` when explicitly authorized (e.g. `--auto_relocate_state`); `SNAPSHOT_ROOT` path must not contain `/workspace/` | `mantis-patch` state-vs-code guard misfires; Block A step 3 confuses SNAPSHOT- vs STATE-relative paths | A:3, D:3; INV-5 | `active_snapshot.root`, `snapshot_root`, `state_root` |
414| 2 | Stale active_snapshot (`active_snapshot.pass != state.pass_number`) | PINNED → STOP or HALT-degrade | Block D step 0: handles same-pass re-entry only; if dir missing → STOP, yield to user | Block A step 2 sentinel may still MATCH (dir retained); CURRENT-PASS CHECK (`active_snapshot.pass == state.pass_number`) required: mismatch → STOP or HALT-degrade (Block B NOT_MATCHED, no authoritative verdicts) | A:2, D:0, B; INV-1, INV-3, INV-4, INV-6 | `active_snapshot.{root, snapshot_id, snapshot_pinned, pass}`, `state.pass_number`, `discovery_commit` |
415| 3 | Pin failure | HALT | Block D step 2/4: skip copy on ENOSPC/error → step 5b; still write `active_snapshot` + pass roots | Authoritative verdicts forbidden; Block B always NOT_MATCHED; reproduce `not_attempted`; patch `VERIFICATION_INCOMPLETE` | D:2, D:4, D:5b; INV-1, INV-2, INV-6 | `active_snapshot.{snapshot_id, snapshot_pinned}` |
416| 4 | Patched shadows | PINNED (pass); `--snapshot_pinned=false` arg | Pass `--target_root=<PATCHED_SHADOW_ROOT>` + `--snapshot_pinned=false` to reattack sub-agent | Block A step 1a: `CODE_ROOT=--target_root` (authoritative); step 2 sentinel SKIPPED (sentinel-EXEMPT) | A:1a, A:2; INV-4 | `target_root`, `snapshot_pinned` (arg), `snapshot_root`, `discovery_commit` |
417| 5 | Different-snapshot duplicate candidates | PINNED | No special action — both passes pinned correctly; dedupe handles it | Block B pairwise: `discovery_commit` differs → NOT_MATCHED → keep ACTIVE + `possible_duplicate_of`; POSSIBLE REGRESSION if archived was RESOLVED | B; INV-3, INV-6 | `discovery_commit`, `possible_duplicate_of`, `status`, `patch_status` |
418| 6 | Absent sink evidence | Any | No special action — Block F is a stage-level mechanical gate | Block F: evidence absent (build error, exit 127, sink unreached) → `not_attempted` (retry-eligible), NEVER `failed_to_reproduce`; HALT ceiling additionally forces `not_attempted` | F; INV-2, INV-6 | `repro_status`, `reattack_status`, `repro_hints` |
419
420**Per-scenario detail:**
421
422**1. Colocated state** (`state_root` nested inside `CODE_ROOT` / snapshot root)
423— The pinned `SNAPSHOT_ROOT` must live under
424`<state_root>/.mantis_snapshots/pass_<N>` (or a clean-VCS worktree/archive), and
425its path **must not** contain the segment `/workspace/` — otherwise
426`mantis-patch`'s state-vs-code path guard misfires (state files appear to be
427"under `CODE_ROOT`"). If `state_root` itself is inside `CODE_ROOT`, the harness
428must HALT-and-yield (safe default) or, when explicitly authorized (e.g.
429`--auto_relocate_state`), relocate it outside the snapshot before pinning. Block
430A step 3 distinguishes SNAPSHOT-RELATIVE path fields (read under `CODE_ROOT`)
431from STATE-RELATIVE fields (read under `state_root/workspace`, never prefixed
432with `CODE_ROOT`); colocation breaks this separation.
433
434**2. Stale active_snapshot** (`active_snapshot.pass != state.pass_number` —
435`active_snapshot` was preserved across the Stage 15 pass increment) — Block D
436step 0 (crash-resume) handles only the SAME-pass re-entry case
437(`active_snapshot.pass == N` → reuse). It does NOT catch a stale snapshot
438carried across the Stage 15 pass increment, because Stage 15 deliberately
439preserves `active_snapshot` while bumping `pass_number` (see Stage 15). Two
440sub-cases:
441
442(a) The prior snapshot dir is now MISSING: Block D step 0 STOPs and yields to
443the user (never re-pin to a possibly-drifted live tree). (b) The prior snapshot
444dir still EXISTS (default keep-2 retention) and its sentinel matches the
445preserved `active_snapshot.snapshot_id`: Block A step 2 sentinel check SUCCEEDS
446(it only compares the sentinel file to `SNAPSHOT_ID`, not to the current pass).
447Block B's pairwise `discovery_commit` check would MATCH a carried-forward
448finding against a new finding stamped with the same stale `SNAPSHOT_ID`,
449silently dropping it as `DUPLICATE` — a false authoritative verdict.
450
451To prevent (b), the HARNESS MUST guarantee that
452`active_snapshot.pass == state.pass_number` before any consumer stage reads it.
453The reference harness (`mantis-meta-agent`) satisfies this by re-pinning every
454pass (Block D step 0 sees `active_snapshot.pass != N` → re-pins → refreshes
455`active_snapshot.pass` before any stage runs), so sub-case (b) never fires
456there. A custom harness that preserves `active_snapshot` across the Stage 15
457pass increment WITHOUT re-pinning MUST either (a) re-pin every pass (the
458reference behavior), or (b) inject an equivalent pre-stage gate that refreshes
459`active_snapshot.pass` or clears `active_snapshot` entirely before invoking
460stages. Stages CANNOT self-detect this staleness via Block B (which is
461`snapshot_id`-only, not `pass`-aware): a carried-forward finding and a new
462finding stamped with the same stale `SNAPSHOT_ID` will MATCH in Block B despite
463the snapshot being stale. The `active_snapshot.pass` field is defined in
464`schema.json` `#/$defs/state/active_snapshot/pass` for exactly this check. The
465harness's Block D step 0 reuse check is NOT a substitute: it only fires on
466same-pass re-entry. (Stages that read `active_snapshot` MAY additionally
467self-check defensively — see each stage's Step 0 sentinel check — but the
468binding guarantee is on the harness.)
469
470**3. Pin failure** (snapshot copy fails — disk full, permissions, too-large
471tree) — Block D step 2 (free-space precheck): compare `du -s` of the live tree
472to `df` free space at `state_root`; if it won't fit → skip copy → step 5b. Block
473D step 4 (failure-tolerant verify): check copy exit status + sanity check (file
474count/size within ~90%); on failure → step 5b (unpinned/HALT). Step 5b:
475`SNAPSHOT_ROOT=<live root>`, `snapshot_pinned=false`,
476`SNAPSHOT_ID="live:"+ISO8601`. The harness still writes `active_snapshot` and
477still passes `--snapshot_root`/`--snapshot_id` to stages so they see the HALT
478signal. Every stage then degrades conservatively: authoritative verdicts
479forbidden (`VERIFIED_SECURE`, `failed_to_reproduce`, `DUPLICATE`,
480`FALSE_POSITIVE`, `NON_VIABLE`); Block B always returns NOT_MATCHED; reproduce
481records `not_attempted`; patch's best attainable is `VERIFICATION_INCOMPLETE`.
482
483**4. Patched shadows** (`--target_root` pointing at a pre-mutated tree;
484sentinel-exempt path 1a in Block A) — `mantis-patch` passes
485`--target_root=<PATCHED_SHADOW_ROOT>` and `--snapshot_pinned=false` to the
486reproduce sub-agent for re-attack verification. Block A step 1a:
487`CODE_ROOT = --target_root` (authoritative override, overrides `--snapshot_root`
488and state fallback). Block A step 2: sentinel check SKIPPED (a `--target_root`
489tree is deliberately mutated and is sentinel-EXEMPT). The
490`--snapshot_pinned=false` argument is the sentinel-exemption, NOT a HALT signal
491— detect HALT by reading STATE (`active_snapshot.snapshot_id` starts with
492`live:`, equivalently `active_snapshot.snapshot_pinned` is `false` in state),
493never from the argument passed on this invocation. The finding's
494`discovery_commit` is unaffected — it retains the pass-level `SNAPSHOT_ID` from
495when it was discovered; only the `--snapshot_pinned=false` argument is local to
496the reattack invocation.
497
498**5. Different-snapshot duplicate candidates** (cross-pass dedupe where
499`discovery_commit` differs — the pairwise Block B NOT_MATCHED path) — Both
500passes pinned correctly; the findings simply come from different snapshots.
501`mantis-dedupe` Block B pairwise check compares the CURRENT finding's
502`discovery_commit` against the ARCHIVED finding's `discovery_commit` (NOT
503against the global `SNAPSHOT_ID`). If they differ → NOT_MATCHED. NOT_MATCHED
504keeps the current finding ACTIVE and sets `possible_duplicate_of` (a soft,
505non-terminal hint — the finding is NOT filtered or trashed). If the archived
506finding was RESOLVED (`patch_status` in {`VERIFIED_SECURE`,
507`MITIGATION_PROPOSED`} OR `status`==`FALSE_POSITIVE` OR
508`production_viability`==`NON_VIABLE`) AND the pair is NOT_MATCHED → POSSIBLE
509REGRESSION: keep ACTIVE, add a history note, never filter (a reverted fix
510re-discovered on new code must never be trashed).
511
512**6. Absent sink evidence** (Block F — PoC compiles but produces no reached-sink
513evidence; `not_attempted` vs `failed_to_reproduce`) — `mantis-reproduce` Block
514F: if EVIDENCE is ABSENT (any compiler/build nonzero exit, exit 127
515command-not-found, exit 2 "No such file", or the sink was never reached) →
516`repro_status = not_attempted` (retry-eligible), STOP. NEVER
517`failed_to_reproduce`. In `--reattack` mode: leave `reattack_status` UNSET with
518a history note "setup_failed" — NEVER `failed_to_bypass`. `failed_to_reproduce`
519is reserved for when the harness PROVABLY reached the vulnerable entrypoint —
520i.e. reached-sink evidence, not setup evidence — but the bug did not fire.
521Reached-sink evidence must originate INSIDE the invoked path or from
522target-produced tracing/backtraces: (a) a PoC script/source harness writes
523`MANTIS_REACHED_ENTRYPOINT` to a sidecar file at the point just before the sink
524call, within its own execution flow (the marker write is part of the invoked
525path, not a pre-launch step); OR (b) for binary/firmware/raw-payload targets,
526the captured crash backtrace or sanitizer trace (ASan/UBSan/MSan/TSan)
527explicitly names the target sink function (target-produced tracing). A marker
528written by an external wrapper BEFORE invoking the target is SETUP EVIDENCE ONLY
529(proves "launch attempted," not "sink reached") and does NOT by itself justify
530`failed_to_reproduce` — treat it as EVIDENCE ABSENT for the decision gate.
531Evidence is recorded in `repro_hints`. In HALT mode, the HALT ceiling
532additionally forces `not_attempted` (no `failed_to_reproduce`), since a negative
533result on an unpinned tree cannot be trusted as authoritative.
534
535______________________________________________________________________
536
537### 6. Semantic Retrieval (RAG) for Large Codebases
538
539For small repositories, the planner can manually scan `workspace/kb/index.md`
540and the researcher can grep for call-sites. At scale (thousands of files, deep
541directory trees, multi-pass campaigns), these approaches miss relevant context
542and waste tokens reading irrelevant files. A semantic retrieval layer lets the
543planner and researcher query for relevant KB entries and code locations without
544reading everything.
545
546Two implementations are supported, sharing the same data contract:
547
548- **Option A (Default — Skill-Based):** A dedicated skill that runs a
549 BM25/TF-IDF helper script over `chunks.jsonl`. Zero external dependencies —
550 works air-gapped, no vector embeddings or vector store required. Optional
551 vector embedding support if available.
552- **Option B (Maximum Scale — MCP-Based):** The harness owns a persistent vector
553 index using vector embeddings, serving persistent `semantic_search_kb` /
554 `semantic_search_code` MCP tools. Better for very large codebases where
555 per-invocation BM25 is too slow.
556
557Both are **opt-in**. The existing skills are not modified; the planner and
558researcher receive runtime instructions to use whichever retrieval mechanism is
559available, falling back to today's manual behavior if neither is present.
560Retrieval results are **coverage HINTs only** — they decide ordering and
561prioritization, never the membership of the audit set. A miss must never cause a
562file, call-site, or investigation to be skipped or dropped.
563
564#### A. Shared Data Contract: `chunks.jsonl`
565
566After Stage 2 (`/mantis-architecture`) completes, chunks are extracted into
567`workspace/kb/chunks.jsonl` (one JSON object per line). The harness can do this
568post-hoc by reading `workspace/kb/*.md`, or the architecture skill can be
569instructed to write it during synthesis as a text-only side effect. Two chunk
570types are produced:
571
5721. **KB chunks** from the existing `workspace/kb/*.md` files:
573
574 ```json
575 {"id": "auth_module:0", "source_file": "workspace/kb/entities/auth_module.md", "entity_type": "entity", "chunk_text": "The auth module handles..."}
576 ```
577
5782. **Code chunks** from `CODE_ROOT` (the pinned snapshot). Each chunk includes
579 the file path and line range so the researcher can request specific files
580 from the snapshot:
581
582 ```json
583 {"id": "src/parser.c:0", "source_file": "src/parser.c", "start_line": 1, "end_line": 80, "chunk_text": "int parse_input(..."}
584 ```
585
586The first line of `chunks.jsonl` is a provenance header recording the
587`SNAPSHOT_ID` the chunks were built against:
588
589```json
590{"_provenance": true, "snapshot_id": "abc123", "kb_snapshot_id": "abc123"}
591```
592
593Before serving queries, check `snapshot_id` in the provenance header against the
594current `SNAPSHOT_ID`; rebuild if they differ. In MODE-OFF (no
595`active_snapshot`), `kb_snapshot_id` is never stamped — skip the index entirely
596and let skills fall back to manual scanning. Never build code chunks from the
597live tree — they must reflect the pinned copy the skills are reading.
598
599#### B. Option A: Skill-Based Retrieval (Default — No Infrastructure)
600
601A dedicated skill reads `chunks.jsonl` and writes+runs a helper script (e.g.
602`workspace/helpers/search_chunks.py`) that performs BM25/TF-IDF similarity
603search. The script is generated by the agent at runtime — no code is shipped
604with the skill (same pattern as `mantis-dedupe`'s `merge_findings.py`). This
605requires zero external dependencies — no embedding model, no vector store, no
606MCP server. It works in air-gapped and VPC-SC environments.
607
608A complete reference blueprint for this skill is available at
609[references/mantis-kb-query.md](references/mantis-kb-query.md). Builders can
610adapt it to their environment. The blueprint includes Block A (Locator
611Resolution), chunk provenance checking, the versioned helper script contract
612(`MANTIS_HELPER_VERSION = 1`), and the JSON output schema.
613
614- **Invocation:** The planner or researcher spawns the skill as a sub-agent with
615 a query string. The skill writes the helper if not already present, runs it,
616 and returns top-K matching chunks as JSON.
617- **Optional embeddings:** If vector embeddings are available, the agent can be
618 instructed to use cosine similarity instead of BM25. This is a runtime
619 configuration toggle, not a different skill.
620- **Snapshot safety:** The skill reads `active_snapshot` from state via Block A
621 (same as every other skill) and checks chunk provenance before serving.
622
623#### C. Option B: MCP-Based Retrieval (For Maximum Scale)
624
625For very large codebases where per-invocation BM25 is too slow, the harness can
626own a persistent vector index using vector embeddings, serving two MCP tools
627(following the same pattern as Guideline 2's Custom MCP for VMs/hardware):
628
629- `semantic_search_kb(query: string) → [{id, source_file, entity_type, chunk_text, score}]`
630 — Searches KB chunks. Returns relevant entity/vulnerability markdown context.
631
632- `semantic_search_code(query: string) → [{file, start_line, end_line, snippet, score}]`
633 — Searches code chunks from the pinned snapshot. Returns relevant code
634 locations.
635
636The harness manages the vector index lifecycle: build from `chunks.jsonl` (or
637directly from `CODE_ROOT`), rebuild when `SNAPSHOT_ID` changes, and handle
638freshness checks. In HALT mode, serve with a `STALE` flag or refuse. In
639MODE-OFF, skip entirely.
640
641#### D. Per-Skill Augmentation Guidance
642
643When a retrieval mechanism (skill or MCP) is available, instruct the following
644skills to use it. These are runtime instructions passed by the harness or
645meta-agent when invoking the skill — the skill files themselves are not
646modified:
647
648- **mantis-architecture:** No changes needed. The harness chunks the existing
649 `workspace/kb/*.md` files after the architect completes Stage 2. If the
650 builder prefers, they may instruct the architect to also write
651 `workspace/kb/chunks.jsonl` during synthesis (Step 3) as a text-only side
652 effect — but this is optional, since the harness can extract chunks post-hoc.
653
654- **mantis-plan:** If a retrieval mechanism is available, instruct the planner
655 to use it to discover `kb_references` for each investigation instead of only
656 manually scanning `workspace/kb/index.md`. For each investigation, query with
657 the investigation title and target file names, then add the top-K matching KB
658 entity/vulnerability files to the `kb_references` array. Manual scanning of
659 `index.md` remains the fallback when no mechanism is available.
660
661- **mantis-researcher:** If a retrieval mechanism is available, instruct Wave 1
662 sub-agents to use it to PRIORITIZE relevant call-sites and cross-module data
663 flows into sinks (e.g., "where does untrusted input reach `memcpy` in the
664 parser module"). Semantic search SUPPLEMENTS grep as a ranking HINT ONLY — it
665 decides ORDER, never MEMBERSHIP
666
667…(truncated)