Multiwoz2.1 Eval

Evaluates a model's ability to track dialogue state (domain, slot, value triplets) across conversation turns, specifically probing its robustness to user mind-changes or 'turnback' utterances that modify previously stated intentions. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint goal accuracy.

qhjqhj00 892fcb0 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiwoz2.1-eval commit 892fcb0df7

Frequently asked questions

npx skillmds add qhjqhj00/multiwoz2-1-eval