Xtragpt Eval

Evaluates an LLM's ability to perform context-aware, instruction-guided revisions of academic paper sections. It probes controllable editing capabilities, specifically measuring adherence to revision instructions, clarity, conciseness, and alignment with scientific writing standards through automated pairwise comparisons and human scoring. Use when the user wants to benchmark on XtraQA, or asks about evaluating this task. Reports Length-controlled (LC) win rate.

qhjqhj00 e58ffba 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/xtragpt-eval commit e58ffbac5f

Frequently asked questions

npx skillmds add qhjqhj00/xtragpt-eval