Artist Eval

This evaluation probes an LLM's ability to perform complex mathematical reasoning and multi-turn function calling by autonomously deciding when and how to invoke external tools. It measures the model's capacity for outcome-based agentic reasoning, including state tracking, error recovery, and precise final answer generation without step-level supervision. Use when the user wants to benchmark on MATH-500, AIME, AMC, Olympiad Bench, BFCL v3, τ-bench, or asks about evaluating this task. Reports Pass@1 accuracy.

qhjqhj00 e6cbf89 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/artist-eval commit e6cbf89435

Frequently asked questions

npx skillmds add qhjqhj00/artist-eval