Character Level F1

This evaluation probes an LLM-based autorater's ability to predict fine-grained machine translation errors (spans, severities, categories) without using human references. It specifically tests how well the model can specialize to a given test set by leveraging in-context examples of human ratings from other systems on the same inputs. Use when the user has predictions and gold and needs to compute character-level F1.

qhjqhj00 c967553 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/character-level-f1 commit c96755341b

Frequently asked questions

npx skillmds add qhjqhj00/character-level-f1