Skilllearnbench Eval

This benchmark evaluates continual learning methods for generating reusable procedural skills in LLM agents. It probes the quality of generated skills, their alignment with execution trajectories, and the ultimate task-solving accuracy and efficiency of a fixed solving agent. Use when the user wants to benchmark on SkillLearnBench, or asks about evaluating this task. Reports Acc..

qhjqhj00 2423290 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/skilllearnbench-eval commit 24232907a3

Frequently asked questions

npx skillmds add qhjqhj00/skilllearnbench-eval