Multiwoz 2.1 Eval

Evaluates a model's ability to track and predict the complete set of user intent slots (dialogue state) across multiple domains in a multi-turn conversation. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports slot accuracy.

qhjqhj00 ef63320 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiwoz-2.1-eval commit ef63320955

Frequently asked questions

npx skillmds add qhjqhj00/multiwoz-2-1-eval