# Nvidia Nemo Gym Gym

> Nemo Gym Reward Profiling

- Skill: `tomevault-io/nvidia-nemo-gym-gym` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add tomevault-io/nvidia-nemo-gym-gym`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tomevault-io/nvidia-nemo-gym-gym/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: tomevault-io (https://skillmd.com/u/tomevault-io)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/tomevault-io/nvidia-nemo-gym-gym

---


# Nemo Gym Reward Profiling

## Invocation Check

Use this skill when the user wants to run, understand, or lightly modify Nemo Gym reward profiling. Keep the answer oriented around the normal workflow:

`gym env start` starts model/resources servers, `gym eval run --no-serve` writes rollout artifacts, and `gym eval profile` generates profiling output from those artifacts.

If the user is primarily debugging a failed job or stack trace, use the `nemo-gym-debugging` skill first.

## Basic Workflow

1. Identify the environment config paths and input JSONL.
2. Start Gym servers with `gym env start`.
3. Collect rollouts with `gym eval run --no-serve`; this writes `rollouts.jsonl` and `*_materialized_inputs.jsonl`.
4. Run `gym eval profile` on the materialized inputs and rollout JSONL to generate `*_reward_profiling.jsonl`.
5. Inspect line counts and profile rows.

Repeated rollouts are the main profiling lever. `num_repeats=1` is valid, but per-task averages and variance are only meaningful with multiple rollouts per task.

## Core Concepts

- `*_materialized_inputs.jsonl`: expanded collection inputs after repeat expansion, agent defaults, and task/rollout id assignment.
- `rollouts.jsonl`: one completed rollout/result per materialized input row.
- `*_reward_profiling.jsonl`: one summarized profile row per original task with at least one completed rollout.
- `_ng_task_index`: original task/sample id.
- `_ng_rollout_index`: repeated rollout id for that task.
- `rollout_infos`: compact per-rollout info inside each task profile row, including reward, token usage, and numeric rollout metrics when available.

Keep reward-to-length or reward-to-token analysis keyed by both `_ng_task_index` and `_ng_rollout_index`.

## Reference Loading

Load references only when the user needs that detail:

- Read `references/quick-start.md` for a generic command template and the minimal run sequence.
- Read `references/output-format.md` to explain materialized inputs, rollout JSONL, reward profile rows, `rollout_infos`, and partial profiling.

## Practical Defaults

- Treat `gym eval profile` as the reward profiling step; rollout collection does not write reward profile files.
- Run strict profiling by default. If rollout collection stopped early, use `++allow_partial_rollouts=True` to profile completed rollouts and drop original input rows with no completed rollout.
- Trust the target checkout's CLI help and `nemo_gym/reward_profile.py` over memory if flags differ.

---
> Source: [NVIDIA-NeMo/Gym](https://github.com/NVIDIA-NeMo/Gym) — distributed by [TomeVault](https://tomevault.io).
<!-- tomevault:4.0:skill_md:2026-06-27 -->

