Autograph R1 Eval

This benchmark evaluates the functional utility of reinforcement learning-optimized knowledge graphs in end-to-end retrieval-augmented generation pipelines. It probes whether task-aware RL training improves both graph-based reasoning and text retrieval performance across multiple question-answering benchmarks and model scales. Use when the user wants to benchmark on Natural Questions (NQ), PopQA, HotpotQA, 2WikiMultihopQA, Musique, or asks about evaluating this task. Reports F1 score.

qhjqhj00 58594ee 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/autograph-r1-eval commit 58594ee3cb

Frequently asked questions

npx skillmds add qhjqhj00/autograph-r1-eval