Eval4nlp 2023 Shared Task Eval

Evaluates reference-free LLM prompting strategies as metrics for machine translation and summarization. It measures how well predicted quality scores correlate with human judgments (MQM for MT, human annotations for summarization). Use when the user wants to benchmark on Eval4NLP 2023 Shared Task (MT & Summarization), or asks about evaluating this task. Reports Kendall correlation.

qhjqhj00 b1d7a92 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/eval4nlp-2023-shared-task-eval commit b1d7a92ad8

Frequently asked questions

npx skillmds add qhjqhj00/eval4nlp-2023-shared-task-eval