Multiwoz Dst Eval

Evaluates a model's ability to track and predict dialogue states across multiple domains in a conversation. It measures how accurately the system maintains slot-value pairs as the user's goals evolve and switches between domains like restaurant, hotel, and taxi. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).

qhjqhj00 80318e0 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiwoz-dst-eval commit 80318e0fba

Frequently asked questions

npx skillmds add qhjqhj00/multiwoz-dst-eval