Statcan Dialogue Eval

Evaluates a model's ability to retrieve relevant statistical data tables from a large corpus based on conversational dialogue history, and its ability to generate appropriate agent responses. It probes intent understanding, table-level grounding, and robustness to temporal distribution shifts. Use when the user wants to benchmark on StatCan Dialogue Dataset, or asks about evaluating this task. Reports recall@10.

qhjqhj00 18f5108 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/statcan-dialogue-eval commit 18f5108b07

Frequently asked questions

npx skillmds add qhjqhj00/statcan-dialogue-eval