Eval Model

Measures the task accuracy of text models served by MAX using standard benchmarks such as GSM8K, MMLU, HellaSwag, ARC, AIME, GPQA, TruthfulQA, WinoGrande, and BABILong. Use when benchmarking a served model, comparing it with model-card or reference scores, verifying that a new MAX model produces correct answers, or running repeatable dataset evaluations against a MAX OpenAI-compatible endpoint.

modular Updated

File contents

modular/skills/tree/main/eval-model commit 6ba8b1f08c

Frequently asked questions

npx skillmds@latest add modular/eval-model