Tombench Eval

Evaluates large language models' Theory of Mind capabilities by testing their ability to infer mental states (beliefs, intentions, emotions) across multiple orders of reasoning using story-based narratives. The benchmark probes whether models can accurately track character perspectives and answer questions about what different agents know or believe in complex social scenarios. Use when the user wants to benchmark on TOMBENCH, or asks about evaluating this task. Reports accuracy.

qhjqhj00 9726f78 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tombench-eval commit 9726f78ae1

Frequently asked questions

npx skillmds add qhjqhj00/tombench-eval