Tid 8 Eval

Probes a model's ability to learn from inherently subjective or disagreed-upon annotations by treating each annotator's label as a separate example, rather than aggregating them into a single ground truth label. Use when the user wants to benchmark on TID-8, or asks about evaluating this task. Reports exact match accuracy.

qhjqhj00 edd4216 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tid-8-eval commit edd42160d4

Frequently asked questions

npx skillmds add qhjqhj00/tid-8-eval