Run Benchmark

Execute a LongMemEval benchmark run - drives per-item hypothesis generation and LLM-as-judge scoring with checkpointing, resumable, capped at 500 items by default

tmuskal de08a35 4.8 KB Updated

File contents

tmuskal/arc-agi-benchmarker/tree/main/plugins/longmemeval-benchmarker/skills/run-benchmark commit de08a3545a

Frequently asked questions

npx skillmds@latest add tmuskal/run-benchmark-2