Ms Tod Eval

Evaluates an LLM agent's ability to retrieve and utilize long-term memory across multiple dialogue sessions to complete goal-oriented tasks. It probes intent-aligned memory selection, slot-level tracking, and dialogue efficiency in maintaining task continuity over extended interactions. Use when the user wants to benchmark on MS-TOD, SGD, MultiWOZ 2.2, or asks about evaluating this task. Reports Success Rate (S.R.).

qhjqhj00 d2bef7a 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ms-tod-eval commit d2bef7a859

Frequently asked questions

npx skillmds add qhjqhj00/ms-tod-eval