Dynamath Eval

Evaluates the robustness of Vision-Language Models in mathematical reasoning by measuring performance across dynamically generated variants of seed questions. It probes how well models handle numerical, geometric, and contextual perturbations while maintaining consistent logical deduction. Use when the user wants to benchmark on DynaMath, or asks about evaluating this task. Reports average-case accuracy.

qhjqhj00 7ed6cae 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dynamath-eval commit 7ed6caeb53

Frequently asked questions

npx skillmds add qhjqhj00/dynamath-eval