Acpbench Hard Eval

Evaluates language and reasoning models on open-ended, generative planning tasks derived from PDDL domains. It probes capabilities like action applicability, reachability, progression, justification, and next-action prediction without predefined answer choices. Use when the user wants to benchmark on ACPBench Hard, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6135ea6 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/acpbench-hard-eval commit 6135ea68bc

Frequently asked questions

npx skillmds add qhjqhj00/acpbench-hard-eval