Aa Omniscience Eval

Evaluates large language models' factual recall and knowledge calibration across domain-specific questions. It measures how reliably models provide correct answers versus hallucinating or abstaining when uncertain, highlighting the gap between raw accuracy and factual reliability. Use when the user wants to benchmark on AA-Omniscience, or asks about evaluating this task. Reports Omniscience Index.

qhjqhj00 86ed93f 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/aa-omniscience-eval commit 86ed93f198

Frequently asked questions

npx skillmds add qhjqhj00/aa-omniscience-eval