Bird Bench Eval

Evaluates the capability of large language models to generate correct SQL queries from natural language questions. It specifically probes how annotation noise and errors in benchmark datasets affect model performance and reliability. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 53e54bf 2.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bird-bench-eval commit 53e54bf397

Frequently asked questions

npx skillmds add qhjqhj00/bird-bench-eval