Tlaps Task Audit

Audit TLAPS Bench proof-completion tasks and cohorts for task integrity, theorem provability evidence, answer leakage, source-reference quality, difficulty signals, stale run artifacts, and model failure causes. Use when reviewing TLAPS task quality, trimming benchmark cohorts, investigating suspicious tasks, analyzing one-shot or agentic results, or preparing a shareable task-audit report.

specula-org ce2c695 11 files · 46.1 KB Updated

File contents

specula-org/tlaps-bench/tree/main/.agents/skills/tlaps-task-audit commit ce2c6952e6

Frequently asked questions

npx skillmds@latest add specula-org/tlaps-task-audit