Human Evaluation Framework

Evaluates text generation models across multiple NLP tasks using standardized human annotation, focusing on reproducibility, annotator quality detection, and scalar scoring of qualities like fluency and correctness. Use when the user has predictions and gold and needs to compute human scores.

qhjqhj00 9e298e3 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/human-evaluation-framework commit 9e298e3179

Frequently asked questions

npx skillmds add qhjqhj00/human-evaluation-framework