Report

Generate a scorecard from a completed LongMemEval run - computes overall accuracy + per-question-type accuracy and writes scorecard.json

tmuskal df98d32 2.0 KB Updated

File contents

tmuskal/arc-agi-benchmarker/tree/main/plugins/longmemeval-benchmarker/skills/report commit df98d3244a

Frequently asked questions

npx skillmds@latest add tmuskal/report-2