Experiment Audit

Audit the experimental **methodology** integrity for a specific claim (Checks A–F: GT provenance, score normalization, result-file existence, dead code, scope, eval-type). Uses cross-model review (external LLM reviewer via llm-chat MCP). The output `overall_verdict` (PASS/WARN/FAIL) is THIS claim's verdict — i.e., whether the claim's experimental process is methodologically sound. Does NOT judge whether the numbers semantically support the claim's hypothesis — that is `/result-to-claim`'s job.

zjunlp 4c9c883 16.6 KB Updated

File contents

zjunlp/mechanist/tree/main/skills/experiment-audit commit 4c9c883a22

Frequently asked questions

npx skillmds@latest add zjunlp/experiment-audit