Kwbench Eval

This benchmark evaluates language models' ability to recognize formal game-theoretic structures (e.g., principal-agent conflict, signaling, strategic omission) in real-world knowledge work scenarios without explicit task hints. It measures the gap between a model's theoretical understanding of these concepts and its capacity for unprompted, practical problem framing. Use when the user wants to benchmark on KWBench, or asks about evaluating this task. Reports Pass Rate.

qhjqhj00 f2e6fec 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/kwbench-eval commit f2e6fec10d

Frequently asked questions

npx skillmds add qhjqhj00/kwbench-eval