Mt Raig Eval

This benchmark evaluates retrieval-augmented insight generation over multiple tables. It requires models to retrieve relevant tables from a database and synthesize multi-hop insights across them, assessing both faithfulness and completeness of the generated reasoning. Use when the user wants to benchmark on MT-RAIG Bench, or asks about evaluating this task. Reports MT-RAIG Eval.

qhjqhj00 49f2222 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mt-raig-eval commit 49f22223c3

Frequently asked questions

npx skillmds add qhjqhj00/mt-raig-eval