Nemo Evaluator Sdk

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking. Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/davila7--claude-code-templates--evaluation-nemo-evaluator commit ecf2eff14d

Frequently asked questions

npx skillmds@latest add tomevault-io/nemo-evaluator-sdk