Cl Gsmsym Eval

Assesses mathematical reasoning and symbolic computation capabilities of LLMs across multiple languages. It uses dynamic, variable-driven templates to generate verifiable ground truths for each instance. The evaluation probes model resilience to linguistic variations and template-specific weaknesses. Use when the user wants to benchmark on CL-GSMSym, or asks about evaluating this task. Reports accuracy.

qhjqhj00 dbe97c8 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cl-gsmsym-eval commit dbe97c8c36

Frequently asked questions

npx skillmds add qhjqhj00/cl-gsmsym-eval