Tabular QA Confidence Eval

Evaluates the calibration and reliability of confidence scores produced by LLMs when answering questions over tabular data. It probes how well predicted confidence aligns with actual accuracy across different elicitation methods and dataset complexities. Use when the user wants to benchmark on WikiTableQuestions, TableBench, or asks about evaluating this task. Reports smooth ECE.

qhjqhj00 23c5847 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tabular-qa-confidence-eval commit 23c584785a

Frequently asked questions

npx skillmds add qhjqhj00/tabular-qa-confidence-eval