Tom Eval

Probes a model's ability to perform first- and second-order Theory of Mind reasoning by tracking agents' true and false beliefs about object locations. It specifically tests whether models can distinguish objective reality from subjective mental states while resisting heuristic shortcuts. Use when the user wants to benchmark on ToMi, or asks about evaluating this task. Reports aggregate accuracy.

qhjqhj00 4e3a065 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tom-eval commit 4e3a065a52

Frequently asked questions

npx skillmds add qhjqhj00/tom-eval