Cti Plausibility Eval

Evaluates whether neural machine translation models correctly rely on contextual cues when generating target tokens. It compares model-extracted cue-target pairs against human-annotated discourse-level expectations to measure the plausibility of context reliance. Use when the user wants to benchmark on SCAT+, or asks about evaluating this task. Reports Macro F1.

qhjqhj00 81cd149 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cti-plausibility-eval commit 81cd149f57

Frequently asked questions

npx skillmds add qhjqhj00/cti-plausibility-eval