Limitgen Eval

Evaluates whether LLMs can accurately identify and articulate critical limitations in scientific research papers across methodological, experimental, analytical, and literature-related dimensions. The benchmark probes the model's ability to ground critiques in domain-specific best practices and produce actionable, substantive feedback rather than superficial presentation critiques. Use when the user wants to benchmark on LimitGen, or asks about evaluating this task. Reports Limitation Quality.

qhjqhj00 e74342b 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/limitgen-eval commit e74342bdbd

Frequently asked questions

npx skillmds add qhjqhj00/limitgen-eval