Symbench Eval

Probes an LLM's ability to solve symbolic reasoning and planning tasks by dynamically switching between textual reasoning and code generation. It evaluates robustness on both seen and unseen tasks, as well as the model's generalizability across different architectures and complexity levels. Use when the user wants to benchmark on SymBench, or asks about evaluating this task. Reports Average Normalized Score (AveNorm).

qhjqhj00 ada00ed 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/symbench-eval commit ada00ed85b

Frequently asked questions

npx skillmds add qhjqhj00/symbench-eval