Dre Bench Eval

Evaluates large language models' fluid intelligence and abstract rule generalization across four hierarchical cognitive levels (Attribute, Spatial, Sequential, Conceptual). It probes the model's ability to dynamically adapt to varying task complexity and apply learned rules to novel grid-based reasoning problems. Use when the user wants to benchmark on DRE-Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 8bd4a83 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dre-bench-eval commit 8bd4a83111

Frequently asked questions

npx skillmds add qhjqhj00/dre-bench-eval