AI Paper Error Audit Eval

Evaluates an LLM-based auditing system's ability to detect, categorize, and quantify objective mistakes in published AI research papers. It measures the system's precision against human verification and its recall against injected ground-truth errors across mathematical, textual, tabular, and cross-reference categories. Use when the user wants to benchmark on Published AI Papers (ICLR, NeurIPS, TMLR), or asks about evaluating this task. Reports precision.

qhjqhj00 8eb1507 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ai-paper-error-audit-eval commit 8eb150718f

Frequently asked questions

npx skillmds add qhjqhj00/ai-paper-error-audit-eval