Confready Checklist Eval

This benchmark evaluates a model's ability to accurately answer conference submission checklist questions based on manuscript content. It specifically probes long-form document understanding, retrieval-augmented generation (RAG) effectiveness, and the model's capacity to reflect on ethical considerations, reproducibility, and societal impacts. Use when the user wants to benchmark on ConfReady Evaluation Set, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 2214f73 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/confready-checklist-eval commit 2214f73288

Frequently asked questions

npx skillmds add qhjqhj00/confready-checklist-eval