Few Shot Bot Eval

Evaluates prompt-based large language models on 15 diverse dialogue tasks, including response generation, conversational parsing, and skill selection, using a few-shot learning setup without fine-tuning. The protocol tests the model's ability to dynamically select the most appropriate task prompt based on dialogue history and generate accurate responses or parses. Use when the user wants to benchmark on Persona Chat, Empathetic Dialogues (ED), Wizard of Wikipedia (WoW), Image Chat (IC), Wizard of Internet (WIT), Controlled Generation (CG-IC), Multi-Session Chat (MSC), DailyDialogue (DD), Stanford Multidomain Dialogue (SMD), DialKG, or asks about evaluating this task. Reports perplexity.

qhjqhj00 f3c6a04 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/few-shot-bot-eval commit f3c6a0486e

Frequently asked questions

npx skillmds add qhjqhj00/few-shot-bot-eval