Mindgames Eval

Evaluates large language models' ability to perform higher-order epistemic reasoning and multi-agent belief tracking. It probes whether models can correctly update beliefs based on public announcements and answer True/False questions about agents' knowledge states. Use when the user wants to benchmark on MindGames, or asks about evaluating this task. Reports accuracy.

qhjqhj00 daccdad 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mindgames-eval commit daccdad4c1

Frequently asked questions

npx skillmds add qhjqhj00/mindgames-eval