QA Translation Fidelity Eval

This benchmark evaluates how well LLM-generated translations preserve the scientific content of original papers. It measures translation fidelity by testing whether a reading model can accurately answer comprehension questions derived from the source text, using only the translated version as context. Use when the user wants to benchmark on Science Across Languages QA Benchmark, or asks about evaluating this task. Reports quiz accuracy.

qhjqhj00 586b1fa 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/qa-translation-fidelity-eval commit 586b1fad4e

Frequently asked questions

npx skillmds add qhjqhj00/qa-translation-fidelity-eval