nemo-evaluator-sdk

orchestra-research/nemo-evaluator-sdk · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄adapter-system.md 9.2 KB
  • 📄configuration.md 9.3 KB
  • 📄custom-benchmarks.md 6.9 KB
  • 📄execution-backends.md 9.1 KB

Related

  1. tao-run-platform · nvidia bundle
    Submit and monitor GPU training jobs on Brev, SLURM, Docker, or Kubernetes using the TAO Execution SDK, with job handles, S3 I/O wrapping, and multi-node distributed training.
    2.2k
    repo stars
  2. tao-run-on-local-docker · nvidia bundle
    Run TAO SDK jobs as Docker containers on a local or remote Docker daemon with NVIDIA GPU support, including preflight checks and credential handling.
    2.2k
    repo stars
  3. tao-finetune-huggingface-model · nvidia bundle
    Fine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
    2.2k
    repo stars
  4. tao-run-on-slurm · nvidia bundle
    Submit and manage TAO training, evaluation, and inference jobs on SLURM GPU clusters over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed storage.
    2.2k
    repo stars
  5. nemo-mbridge-multi-node-slurm · nvidia bundle
    Convert single-node PyTorch distributed scripts into multi-node Slurm sbatch jobs and debug common multi-node failures, covering srun-native and torch.distributed approaches, container setup, NCCL timeouts, and interactive allocation.
    2.2k
    repo stars
  6. rag-perf · nvidia bundle
    Run config-driven performance benchmarks against a deployed NVIDIA RAG Blueprint server, including profiling and load testing, with a unified report.
    2.2k
    repo stars

Frequently asked questions

How do I install the nemo-evaluator-sdk skill?

Run npx skillmds add orchestra-research/nemo-evaluator-sdk in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the nemo-evaluator-sdk skill do?

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms. It is listed under AI & ML, DevOps & Infra, Containers & Kubernetes, Model Training & Fine-tuning on SkillMD.

Is nemo-evaluator-sdk safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with nemo-evaluator-sdk?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is nemo-evaluator-sdk free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published nemo-evaluator-sdk?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.