Lllm Paper Filtering Eval

Evaluates the ability of LLMs to accurately classify academic papers as discussing LLM limitations and to extract supporting evidence from abstracts. It measures alignment with human expert annotations using ordinal rating agreement and span-level extraction metrics. Use when the user wants to benchmark on ACL Anthology & arXiv (crawled 2022-2025), or asks about evaluating this task. Reports weighted-cohens-kappa.

qhjqhj00 504a92a 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/lllm-paper-filtering-eval commit 504a92a31d

Frequently asked questions

npx skillmds add qhjqhj00/lllm-paper-filtering-eval