Multiwoz Eval

Evaluates end-to-end task-oriented dialogue systems on their ability to track user goals, fulfill multi-domain requests, and generate contextually appropriate responses. It specifically probes how well models maintain conversation state and achieve user objectives without relying on full historical dialogue context. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports inform rate, success rate.

qhjqhj00 ce7e392 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiwoz-eval commit ce7e392944

Frequently asked questions

npx skillmds add qhjqhj00/multiwoz-eval