Rl Dialogue Benchmark Eval

Evaluates the robustness and generalization of reinforcement learning-based dialogue management policies across varying simulated environments. It probes how well RL algorithms handle different domain sizes, user behavior profiles, and noisy speech input channels in task-oriented spoken dialogue systems. Use when the user wants to benchmark on PyDial simulated environments, or asks about evaluating this task. Reports average success rate.

qhjqhj00 c853c4b 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/rl-dialogue-benchmark-eval commit c853c4b141

Frequently asked questions

npx skillmds add qhjqhj00/rl-dialogue-benchmark-eval