Harbor F2p P2p Deep Dive

Deep dive on Harbor trial results for tasks that use SWE-Bench-style F2P (FAIL_TO_PASS) and P2P (PASS_TO_PASS) reference tests. Diagnoses why an agent failed and audits whether a failing task is genuinely hard or unfair (instruction-vs-verifier mismatch). Use when the user asks to analyze, debug, investigate, or deep dive on a Harbor run, trial, or job directory; when the user says "harbor deep dive", "f2p deep dive", or "p2p deep dive"; when a trial scores lower than expected; or when the user wants to assess task fairness, instruction quality, verifier correctness, or whether an F2P or P2P test is reasonable for the instruction given. Works with any Harbor task whose verifier reports F2P/P2P-style results, and with any supported agent (claude-code, opencode, codex). Produces a per-failure root-cause verdict (capability gap / verifier issue / instruction issue) by triangulating verifier output, the agent's implementation, the reference tests, and the instruction.

NVIDIA NeMo Updated

File contents

nvidia-nemo/switchyard/tree/main/experimental/craft-taskgen/.claude/skills/harbor-f2p-p2p-deep-dive commit 1e9ac1b331

Frequently asked questions

npx skillmds@latest add nvidia-nemo/harbor-f2p-p2p-deep-dive