Simpleqa Verified Eval

This benchmark evaluates an LLM's parametric factuality and internal knowledge recall on short-form questions. It measures whether models can correctly answer factual queries without relying on external search tools or retrieval augmentations. Use when the user wants to benchmark on SimpleQA Verified, or asks about evaluating this task. Reports F1-Score.

qhjqhj00 3be72d3 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/simpleqa-verified-eval commit 3be72d34cf

Frequently asked questions

npx skillmds add qhjqhj00/simpleqa-verified-eval