Truthfulqa Eval

Evaluates the factual accuracy and truthfulness of large language models by measuring their ability to select correct answers over common misconceptions. It probes the model's capacity to resist generating plausible but false statements across diverse categories like health, law, and politics. The benchmark specifically tests whether models can identify and output factually correct responses when presented with multiple candidate answers. Use when the user wants to benchmark on TruthfulQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 47fb5c2 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/truthfulqa-eval commit 47fb5c2234

Frequently asked questions

npx skillmds add qhjqhj00/truthfulqa-eval