Cmu Doq Eval

Evaluates the quality of generated responses in document-grounded conversations, specifically measuring how well models leverage external document context to produce engaging and fluent multi-turn dialogue. It assesses both automatic language modeling metrics and human-perceived response quality. Use when the user wants to benchmark on CMU.DoG, or asks about evaluating this task. Reports Perplexity.

qhjqhj00 e6071c3 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cmu-doq-eval commit e6071c3721

Frequently asked questions

npx skillmds add qhjqhj00/cmu-doq-eval