NVIDIA DiffusionHarmonizer (NuRec post-processing)
Purpose
Run NVIDIA DiffusionHarmonizer on rendered images from neural
reconstructions. DiffusionHarmonizer is a single-step,
temporally-aware image diffusion enhancer for NeRF / 3DGS /
NuRec-style renderings. It improves realism, reduces
reconstruction artifacts, and harmonizes inserted dynamic
objects with the surrounding scene.
When to Use / When NOT to Use
Use this skill when the user has rendered frames from
NuRec, 3DGS, NeRF, or a similar reconstruction pipeline and
wants to enhance, harmonize, evaluate, or optionally fine-tune
the DiffusionHarmonizer model.
Do NOT use this skill when:
- The user wants to train or render the 3D reconstruction itself
(use the
nurec skills).
- The user wants to convert raw sensor data to NCore V4 (use
ncore).
- The user wants a generic photo enhancer. DiffusionHarmonizer
is tuned for neural-reconstruction artifacts and
object-insertion failures.
- The user only wants NuRec inline rendering with
--enable-difix. That remains a NuRec runtime feature; use the
nurec skills for the complete serve-grpc / render-grpc
command shape.
What changed from the older Fixer skill
This skill follows the public NVIDIA/harmonizer release, not
the older NGC JIT .pt artifact recipe. Use these public
release artifacts:
- Code: https://github.com/NVIDIA/harmonizer
- Model:
nvidia/Harmonizer on Hugging Face (the paper checkpoint
models/diffusion_harmonizer.pkl), plus the base
nvidia/Cosmos-Predict2-0.6B-Text2Image model that inference
also requires.
- Checkpoint download:
./download_checkpoints.sh from the repo
root. It fetches the Harmonizer checkpoints into models/
(diffusion_harmonizer.pkl, harmonizer_nontemporal.pt) and
the base Cosmos DiT + tokenizer into
src/checkpoints/nvidia/Cosmos-Predict2-0.6B-Text2Image/.
- Runtime: the
harmonizer-cosmos-env image built from
Dockerfile.cosmos (base nvcr.io/nvidia/pytorch:25.10-py3).
- Inference entry:
src/inference_pix2pix_turbo_harmonizer.py,
run from inside /work/src so it can import its sibling
modules.
- Evaluation entry:
src/evaluate_test_dataset.py
- Training entry:
src/train_pix2pix_turbo_harmonizer.py
Do not run inference_pretrained_model.py; the current README
documents inference_pix2pix_turbo_harmonizer.py as the
inference entry point.
Background
DiffusionHarmonizer is described in DiffusionHarmonizer:
Bridging Neural Reconstruction and Photorealistic Simulation
with Online Diffusion Enhancer
(arXiv 2602.24096, CVPR 2026).
It distills a pretrained multi-step diffusion model into a
single-step enhancer designed for online simulation and offline
data cleanup.
Two operating modes:
- Offline: clean pseudo-training views rendered from a
reconstruction, then distill the improved views back into the
3D representation.
- Online: enhance frames during simulation/inference by
harmonizing color and lighting, reconstructing missing or
inconsistent shadows for inserted actors, and reducing residual
reconstruction artifacts.
The public model card describes DiffusionHarmonizer-cosmos-0.6B,
a Cosmos Predict2 Diffusion Transformer post-trained at
576x1024 input and output resolution.
Inputs
- code_dir — checkout of
https://github.com/NVIDIA/harmonizer.
- model_dir — checkpoints fetched by
./download_checkpoints.sh
from the repo root. It places the paper checkpoint at
models/diffusion_harmonizer.pkl and the required base Cosmos
model under
src/checkpoints/nvidia/Cosmos-Predict2-0.6B-Text2Image/.
- input_dir — directory of rendered RGB frames (
.png,
.jpg, .jpeg) to enhance.
- output_dir —
inference_pix2pix_turbo_harmonizer.py does
not take an output flag; it writes to a sibling folder named
<input_dir>_<model_identifier> next to the input directory.
- HF_TOKEN — Hugging Face token. Not required for
nvidia/Harmonizer, which is public and downloads anonymously.
Required, with licence acceptance, for the gated
nvidia/Cosmos-Predict2-0.6B-Text2Image and (if used)
nvidia/Harmonizer-Dataset. A cached hf auth login works in place
of the environment variable.
- NGC_API_KEY — often needed to authenticate
docker pull
from nvcr.io. Use only for container pulls, not model
download.
Instructions
Validate the host. Have the agent execute
this skill's scripts/validate_setup.py via its standard script
runner — e.g. run_script("scripts/validate_setup.py") or, from this
skill's own directory, python scripts/validate_setup.py. Paths
beginning scripts/ in this document are relative to the skill
directory; this repository has no scripts/ at its root, so the
command will not resolve from the checkout. It checks Docker, the
NVIDIA Container Toolkit, GPU architecture, git, the
Hugging Face CLI, token presence, and free disk space. It exits
non-zero for missing Docker, toolkit, GPU or git. A missing
Hugging Face CLI or token is only a warning that still exits
zero — pass --strict to make warnings fail. Two further gaps to
know: it accepts the legacy huggingface-cli, but
download_checkpoints.sh invokes hf only, so that combination
passes the check and then fails at download; and its disk check
covers inference only, not the ~1.76 TB dataset.
Clone the code and build (or pull) the runtime image.
Full commands and the Blackwell patch caveat live in
references/inference.md.
If you are already inside a harmonizer checkout — which is the case when
this skill ships in-repo — use it and skip the clone; cloning from
there nests a second checkout and switches work to upstream main.
# only when you do not already have a checkout
git clone https://github.com/NVIDIA/harmonizer.git
cd harmonizer
# from an existing checkout, start here
docker build -t harmonizer-cosmos-env -f Dockerfile.cosmos .
Download the checkpoints. From the repo root run the
helper, which fetches both the Harmonizer checkpoints and the
base Cosmos model into the paths the code expects.
export HF_TOKEN="hf_your_token_here" # your real token, no angle brackets
hf auth login --token "$HF_TOKEN"
./download_checkpoints.sh
Verify models/diffusion_harmonizer.pkl and
src/checkpoints/nvidia/Cosmos-Predict2-0.6B-Text2Image/
exist.
Confirm input_dir exists and filenames sort into frame
order. The temporal inference script sorts frames with
plain lexical sort (not natural sort, despite the video-assembly step
using natsorted) and uses previous outputs as references, so
fixed-width zero-padded filenames are required — frame_10.png
otherwise sorts before frame_2.png and corrupts the sequence; prefer
zero-padded names such as frame_000001.png.
Run inference inside the container with the repo mounted
at /work, then cd /work/src and run
inference_pix2pix_turbo_harmonizer.py. Pin
-u $(id -u):$(id -g) so outputs are owned by the host user.
Output frames land in
<input_dir>_<model_identifier>. Full docker run recipe and
flag matrix in
references/inference.md.
Validate that the output frame count matches the input
frame count, spot-check frames, and (if ground truth is
available) consult
references/evaluation.md for the paired
PSNR/LPIPS shape. Do not promise the user that evaluation will run —
evaluate_test_dataset.py has upstream blockers documented in that file.
(Optional) Train or fine-tune. Download the dataset (or
prepare JSON manifests in the documented format), then run
src/train_pix2pix_turbo_harmonizer.py with the recommended
hyperparameters. For fine-tuning, initialize from the released
checkpoint with --pretrained_path /path/to/diffusion_harmonizer.pkl. Full recipe + NuRec
data-pair recipes in
references/training.md.
(Optional) Teardown. Follow
references/teardown.md to remove
images, code clones, model weights, datasets, and outputs.
Examples
Example 1 — Enhance a folder of rendered frames
Run scripts/validate_setup.py, build the image once (see
references/wrapper-image.md), then
invoke inference_pix2pix_turbo_harmonizer.py inside the
harmonizer-cosmos-env container with the repo checkout mounted at
/work. From /work/src, point --input_image at the rendered
frames, --model_path at /work/models/diffusion_harmonizer.pkl,
set --model_identifier, and pass typical flags
--timestep 250 --resolution 1024 --use_sched. Enhanced frames are
written to <input_dir>_<model_identifier>. The canonical
docker run command and the full flag matrix live in
references/inference.md.
Example 2 — Quantitative PSNR / LPIPS evaluation
Prepare the paired test_dataset/{scene}/render +
test_dataset/{scene}/gt layout. src/evaluate_test_dataset.py is the
intended entry point but does not currently run as shipped — see
references/evaluation.md for the exact
directory shape, the docker run shape, and the two upstream blockers.
Example 3 — Fine-tune from the public checkpoint
Download nvidia/Harmonizer-Dataset, prepare the
training JSON, then run
src/train_pix2pix_turbo_harmonizer.py with the multi-GPU
accelerate launch recipe in
references/training.md. For
fine-tuning add --pretrained_path /path/to/diffusion_harmonizer.pkl. Do not add --fixing_data_weight
here: with the single /data/data.json source shown above it cannot do
anything, because every item in a source gets the same weight and the
multiplier normalises out
(src/train_pix2pix_turbo_harmonizer.py:200-208). Up-weighting needs a
comma-separated --dataset_folder with the correction data isolated in its
own source whose path contains nre_data, plus --weighted_sampler.
Example 4 — Non-temporal (frame-by-frame) enhancement
inference_pix2pix_turbo_harmonizer.py is temporal by default
and uses previous enhanced frames as references
(--offset_list -1 -2 -3 -4). When the user wants each frame
enhanced independently (e.g. unordered images), add
--nontemporal to disable temporal conditioning. Full command +
--offset_list defaults in
references/inference.md.
Prerequisites
- OS: Linux host.
- GPU / driver: NVIDIA GPU Ampere or newer (compute
capability
>= 8.0; A100, A10, L40, H100, RTX 30/40/PRO,
B200, GB200).
- Container runtime: Docker with the NVIDIA Container
Toolkit.
- Tools:
git, python3, Hugging Face CLI (hf or
huggingface-cli).
- Secrets:
HF_TOKEN with the gated
nvidia/Cosmos-Predict2-0.6B-Text2Image
licenses accepted (required to download model weights and
the optional dataset).
NGC_API_KEY (often required for docker login nvcr.io
before pulling nvcr.io/nvidia/pytorch:25.10-py3).
- Disk: at least
120 GB free for the runtime image, build cache,
model weights and outputs — this covers inference only. The optional
training dataset is a further **1.76 TB**, not included in that figure
nor checked by the prerequisite script; see
references/training.md before downloading it.
- Source / model / dataset:
scripts/validate_setup.py checks the above. It fails hard on missing
Docker, NVIDIA Container Toolkit, a supported GPU or git; Hugging Face
CLI and token issues are warnings unless you pass --strict. It does not
verify space for the optional dataset.
Scripts
| Script |
Purpose |
Usage |
scripts/validate_setup.py |
Verify Docker, NVIDIA Container Toolkit, GPU architecture, git, Hugging Face CLI, token presence, and disk space. No network calls. |
run_script("scripts/validate_setup.py") or python scripts/validate_setup.py |
scripts/.env.example |
Template for HF_TOKEN and optional NGC_API_KEY. Keep the filled-in file OUTSIDE the checkout. Dockerfile.cosmos:29 is COPY . /harmonizer-codebase, and this repository ships no .dockerignore, so anything in the checkout — including a .env — is uploaded to the Docker daemon or remote builder and retained in build cache. A .gitignore does not scope build context and will not prevent this. |
install -m 600 scripts/.env.example ~/.harmonizer.env && set -a && . ~/.harmonizer.env && set +a (the template is mode 0664 — a plain cp would leave your live tokens readable by other local users) |
References
references/inference.md — container
build, raw-base fallback, Blackwell patches, checkpoint
download, inference_pix2pix_turbo_harmonizer.py flag matrix,
non-temporal mode.
references/evaluation.md — paired
test_dataset/ layout and evaluate_test_dataset.py command, plus the
upstream blockers that stop it running today.
references/training.md — dataset
download, training JSON format, multi-GPU accelerate launch
recipe, fine-tuning flags, NuRec data-pair recipes.
references/wrapper-image.md —
build and run the project image for repeat inference.
references/troubleshooting.md
— extended diagnostic notes.
references/teardown.md — cleanup
inventory for images, code, Hugging Face caches, datasets,
and outputs.
- Public code: https://github.com/NVIDIA/harmonizer
- Model card: https://huggingface.co/nvidia/Harmonizer
- Dataset: https://huggingface.co/datasets/nvidia/Harmonizer-Dataset
- Paper: https://arxiv.org/abs/2602.24096
- Project page: https://research.nvidia.com/labs/sil/projects/diffusion-harmonizer/
Limitations
- Rendered inputs only. The model is tuned for
neural-reconstruction renderings and object-insertion
artifacts, not arbitrary real photos.
- Primary model resolution is 576x1024. The inference script
maps resolution key
1024 to 1024x576. Only 1024, 960,
and 1360 are supported keys; 1024 matches the model-card
operating point.
- Temporal references and filename order matter. The
inference script enhances frames in plain lexically sorted order
(zero-pad filenames to a fixed width) and
feeds previous outputs back as temporal references. Use
zero-padded frame numbers, or pass
--nontemporal for
unordered images.
- Container builds can be large. The runtime image, build
cache, model weights, and optional dataset can exceed 100 GB.
- Training is multi-GPU by default. The README command
assumes 8 GPUs with bf16 mixed precision.
- Public code evolves. The current README documents
inference_pix2pix_turbo_harmonizer.py as the inference entry
point; if a future checkout renames it, prefer the script and
flags present in that checkout.
Troubleshooting (top 5)
| Error / symptom |
Most common cause |
docker: could not select device driver ... gpu |
NVIDIA Container Toolkit missing or Docker is not configured for the NVIDIA runtime. |
docker pull 401 / 403 from nvcr.io |
Docker is not authenticated to NGC, or the API key lacks container access. |
hf download ... 401 / 403 |
Hitting a gated repo (nvidia/Cosmos-Predict2-0.6B-Text2Image, nvidia/Harmonizer-Dataset) without an accepted licence or a valid token. nvidia/Harmonizer itself is public and needs neither. |
diffusion_harmonizer.pkl missing |
Checkpoint download path is wrong or incomplete. Re-run ./download_checkpoints.sh from the repo root. |
Output files owned by root |
The docker run omitted -u $(id -u):$(id -g). |
Full matrix in
references/troubleshooting.md.
Teardown
A full workflow can leave large artifacts on disk: the Cosmos
image, project image, build cache, harmonizer code checkout,
Hugging Face model weights, optional dataset, evaluation
outputs, and enhanced frames. Reclaim them with the inventory
in references/teardown.md. Do not
revoke HF_TOKEN or NGC_API_KEY as normal cleanup. Rotate a
token only if you suspect it was leaked.
1---2name: harmonizer3description: Use to run NVIDIA DiffusionHarmonizer (public successor to the older Fixer recipes) to enhance, harmonize, evaluate, or fine-tune novel-view frames from NuRec / 3DGS / NeRF reconstructions. Do NOT use for training the 3D reconstruction itself (use the `nurec` skills) or for sensor-to-NCore conversion (use `ncore`).4license: CC-BY-4.0 AND Apache-2.05---67# NVIDIA DiffusionHarmonizer (NuRec post-processing)89## Purpose1011Run NVIDIA DiffusionHarmonizer on rendered images from neural12reconstructions. DiffusionHarmonizer is a single-step,13temporally-aware image diffusion enhancer for NeRF / 3DGS /14NuRec-style renderings. It improves realism, reduces15reconstruction artifacts, and harmonizes inserted dynamic16objects with the surrounding scene.1718## When to Use / When NOT to Use1920**Use this skill when** the user has rendered frames from21NuRec, 3DGS, NeRF, or a similar reconstruction pipeline and22wants to enhance, harmonize, evaluate, or optionally fine-tune23the DiffusionHarmonizer model.2425**Do NOT use this skill when:**2627- The user wants to train or render the 3D reconstruction itself28 (use the `nurec` skills).29- The user wants to convert raw sensor data to NCore V4 (use30 `ncore`).31- The user wants a generic photo enhancer. DiffusionHarmonizer32 is tuned for neural-reconstruction artifacts and33 object-insertion failures.34- The user only wants NuRec inline rendering with35 `--enable-difix`. That remains a NuRec runtime feature; use the36 `nurec` skills for the complete `serve-grpc` / `render-grpc`37 command shape.3839## What changed from the older Fixer skill4041This skill follows the public `NVIDIA/harmonizer` release, not42the older NGC JIT `.pt` artifact recipe. Use these public43release artifacts:4445- Code: <https://github.com/NVIDIA/harmonizer>46- Model: `nvidia/Harmonizer` on Hugging Face (the paper checkpoint47 `models/diffusion_harmonizer.pkl`), plus the base48 `nvidia/Cosmos-Predict2-0.6B-Text2Image` model that inference49 also requires.50- Checkpoint download: `./download_checkpoints.sh` from the repo51 root. It fetches the Harmonizer checkpoints into `models/`52 (`diffusion_harmonizer.pkl`, `harmonizer_nontemporal.pt`) and53 the base Cosmos DiT + tokenizer into54 `src/checkpoints/nvidia/Cosmos-Predict2-0.6B-Text2Image/`.55- Runtime: the `harmonizer-cosmos-env` image built from56 `Dockerfile.cosmos` (base `nvcr.io/nvidia/pytorch:25.10-py3`).57- Inference entry: `src/inference_pix2pix_turbo_harmonizer.py`,58 run from inside `/work/src` so it can import its sibling59 modules.60- Evaluation entry: `src/evaluate_test_dataset.py`61- Training entry: `src/train_pix2pix_turbo_harmonizer.py`6263Do not run `inference_pretrained_model.py`; the current README64documents `inference_pix2pix_turbo_harmonizer.py` as the65inference entry point.6667## Background6869DiffusionHarmonizer is described in *DiffusionHarmonizer:70Bridging Neural Reconstruction and Photorealistic Simulation71with Online Diffusion Enhancer*72([arXiv 2602.24096](https://arxiv.org/abs/2602.24096), CVPR 2026).73It distills a pretrained multi-step diffusion model into a74single-step enhancer designed for online simulation and offline75data cleanup.7677Two operating modes:7879- **Offline:** clean pseudo-training views rendered from a80 reconstruction, then distill the improved views back into the81 3D representation.82- **Online:** enhance frames during simulation/inference by83 harmonizing color and lighting, reconstructing missing or84 inconsistent shadows for inserted actors, and reducing residual85 reconstruction artifacts.8687The public model card describes `DiffusionHarmonizer-cosmos-0.6B`,88a Cosmos Predict2 Diffusion Transformer post-trained at89`576x1024` input and output resolution.9091## Inputs9293- **code_dir** — checkout of94 <https://github.com/NVIDIA/harmonizer>.95- **model_dir** — checkpoints fetched by `./download_checkpoints.sh`96 from the repo root. It places the paper checkpoint at97 `models/diffusion_harmonizer.pkl` and the required base Cosmos98 model under99 `src/checkpoints/nvidia/Cosmos-Predict2-0.6B-Text2Image/`.100- **input_dir** — directory of rendered RGB frames (`.png`,101 `.jpg`, `.jpeg`) to enhance.102- **output_dir** — `inference_pix2pix_turbo_harmonizer.py` does103 not take an output flag; it writes to a sibling folder named104 `<input_dir>_<model_identifier>` next to the input directory.105- **HF_TOKEN** — Hugging Face token. **Not** required for106 `nvidia/Harmonizer`, which is public and downloads anonymously.107 Required, with licence acceptance, for the gated108 `nvidia/Cosmos-Predict2-0.6B-Text2Image` and (if used)109 `nvidia/Harmonizer-Dataset`. A cached `hf auth login` works in place110 of the environment variable.111- **NGC_API_KEY** — often needed to authenticate `docker pull`112 from `nvcr.io`. Use only for container pulls, not model113 download.114115## Instructions1161171. **Validate the host.** Have the agent execute118 this skill's `scripts/validate_setup.py` via its standard script119 runner — e.g. `run_script("scripts/validate_setup.py")` or, from this120 skill's own directory, `python scripts/validate_setup.py`. Paths121 beginning `scripts/` in this document are relative to the skill122 directory; this repository has no `scripts/` at its root, so the123 command will not resolve from the checkout. It checks Docker, the124 NVIDIA Container Toolkit, GPU architecture, `git`, the125 Hugging Face CLI, token presence, and free disk space. It exits126 non-zero for missing Docker, toolkit, GPU or `git`. A missing127 Hugging Face CLI or token is only a **warning** that still exits128 zero — pass `--strict` to make warnings fail. Two further gaps to129 know: it accepts the legacy `huggingface-cli`, but130 `download_checkpoints.sh` invokes `hf` only, so that combination131 passes the check and then fails at download; and its disk check132 covers inference only, not the ~1.76 TB dataset.1332. **Clone the code and build (or pull) the runtime image.**134 Full commands and the Blackwell patch caveat live in135 [`references/inference.md`](references/inference.md).136137 If you are already inside a harmonizer checkout — which is the case when138 this skill ships in-repo — **use it and skip the clone**; cloning from139 there nests a second checkout and switches work to upstream `main`.140141 ```bash142 # only when you do not already have a checkout143 git clone https://github.com/NVIDIA/harmonizer.git144 cd harmonizer145146 # from an existing checkout, start here147 docker build -t harmonizer-cosmos-env -f Dockerfile.cosmos .148 ```1491503. **Download the checkpoints.** From the repo root run the151 helper, which fetches both the Harmonizer checkpoints and the152 base Cosmos model into the paths the code expects.153154 ```bash155 export HF_TOKEN="hf_your_token_here" # your real token, no angle brackets156 hf auth login --token "$HF_TOKEN"157 ./download_checkpoints.sh158 ```159160 Verify `models/diffusion_harmonizer.pkl` and161 `src/checkpoints/nvidia/Cosmos-Predict2-0.6B-Text2Image/`162 exist.1634. **Confirm `input_dir` exists and filenames sort into frame164 order.** The temporal inference script sorts frames with165 plain lexical sort (not natural sort, despite the video-assembly step166 using `natsorted`) and uses previous outputs as references, so167 fixed-width zero-padded filenames are **required** — `frame_10.png`168 otherwise sorts before `frame_2.png` and corrupts the sequence; prefer169 zero-padded names such as `frame_000001.png`.1705. **Run inference inside the container** with the repo mounted171 at `/work`, then `cd /work/src` and run172 `inference_pix2pix_turbo_harmonizer.py`. Pin173 `-u $(id -u):$(id -g)` so outputs are owned by the host user.174 Output frames land in175 `<input_dir>_<model_identifier>`. Full `docker run` recipe and176 flag matrix in177 [`references/inference.md`](references/inference.md).1786. **Validate that the output frame count matches the input179 frame count,** spot-check frames, and (if ground truth is180 available) consult181 [`references/evaluation.md`](references/evaluation.md) for the paired182 PSNR/LPIPS shape. **Do not promise the user that evaluation will run** —183 `evaluate_test_dataset.py` has upstream blockers documented in that file.1847. **(Optional) Train or fine-tune.** Download the dataset (or185 prepare JSON manifests in the documented format), then run186 `src/train_pix2pix_turbo_harmonizer.py` with the recommended187 hyperparameters. For fine-tuning, initialize from the released188 checkpoint with `--pretrained_path189 /path/to/diffusion_harmonizer.pkl`. Full recipe + NuRec190 data-pair recipes in191 [`references/training.md`](references/training.md).1928. **(Optional) Teardown.** Follow193 [`references/teardown.md`](references/teardown.md) to remove194 images, code clones, model weights, datasets, and outputs.195196## Examples197198### Example 1 — Enhance a folder of rendered frames199200Run `scripts/validate_setup.py`, build the image once (see201[`references/wrapper-image.md`](references/wrapper-image.md)), then202invoke `inference_pix2pix_turbo_harmonizer.py` inside the203`harmonizer-cosmos-env` container with the repo checkout mounted at204`/work`. From `/work/src`, point `--input_image` at the rendered205frames, `--model_path` at `/work/models/diffusion_harmonizer.pkl`,206set `--model_identifier`, and pass typical flags207`--timestep 250 --resolution 1024 --use_sched`. Enhanced frames are208written to `<input_dir>_<model_identifier>`. The canonical209`docker run` command and the full flag matrix live in210[`references/inference.md`](references/inference.md).211212### Example 2 — Quantitative PSNR / LPIPS evaluation213214Prepare the paired `test_dataset/{scene}/render` +215`test_dataset/{scene}/gt` layout. `src/evaluate_test_dataset.py` is the216intended entry point but does not currently run as shipped — see217[`references/evaluation.md`](references/evaluation.md) for the exact218directory shape, the `docker run` shape, and the two upstream blockers.219220### Example 3 — Fine-tune from the public checkpoint221222Download `nvidia/Harmonizer-Dataset`, prepare the223training JSON, then run224`src/train_pix2pix_turbo_harmonizer.py` with the multi-GPU225`accelerate launch` recipe in226[`references/training.md`](references/training.md). For227fine-tuning add `--pretrained_path228/path/to/diffusion_harmonizer.pkl`. Do **not** add `--fixing_data_weight`229here: with the single `/data/data.json` source shown above it cannot do230anything, because every item in a source gets the same weight and the231multiplier normalises out232(`src/train_pix2pix_turbo_harmonizer.py:200-208`). Up-weighting needs a233comma-separated `--dataset_folder` with the correction data isolated in its234own source whose path contains `nre_data`, plus `--weighted_sampler`.235236### Example 4 — Non-temporal (frame-by-frame) enhancement237238`inference_pix2pix_turbo_harmonizer.py` is temporal by default239and uses previous enhanced frames as references240(`--offset_list -1 -2 -3 -4`). When the user wants each frame241enhanced independently (e.g. unordered images), add242`--nontemporal` to disable temporal conditioning. Full command +243`--offset_list` defaults in244[`references/inference.md`](references/inference.md).245246## Prerequisites247248- **OS:** Linux host.249- **GPU / driver:** NVIDIA GPU Ampere or newer (compute250 capability `>= 8.0`; A100, A10, L40, H100, RTX 30/40/PRO,251 B200, GB200).252- **Container runtime:** Docker with the NVIDIA Container253 Toolkit.254- **Tools:** `git`, `python3`, Hugging Face CLI (`hf` or255 `huggingface-cli`).256- **Secrets:**257 - `HF_TOKEN` with the gated258 [`nvidia/Cosmos-Predict2-0.6B-Text2Image`](https://huggingface.co/nvidia/Cosmos-Predict2-0.6B-Text2Image)259 licenses accepted (required to download model weights and260 the optional dataset).261 - `NGC_API_KEY` (often required for `docker login nvcr.io`262 before pulling `nvcr.io/nvidia/pytorch:25.10-py3`).263- **Disk:** at least ~120 GB free for the runtime image, build cache,264 model weights and outputs — this covers **inference only**. The optional265 training dataset is a further **~1.76 TB**, not included in that figure266 nor checked by the prerequisite script; see267 [`references/training.md`](references/training.md) before downloading it.268- **Source / model / dataset:**269 - Code: <https://github.com/NVIDIA/harmonizer>.270 - Model: <https://huggingface.co/nvidia/Harmonizer>.271 - Base model:272 <https://huggingface.co/nvidia/Cosmos-Predict2-0.6B-Text2Image>.273 - Optional dataset:274 <https://huggingface.co/datasets/nvidia/Harmonizer-Dataset>.275276`scripts/validate_setup.py` checks the above. It fails hard on missing277Docker, NVIDIA Container Toolkit, a supported GPU or `git`; Hugging Face278CLI and token issues are warnings unless you pass `--strict`. It does not279verify space for the optional dataset.280281## Scripts282283| Script | Purpose | Usage |284|--------|---------|-------|285| `scripts/validate_setup.py` | Verify Docker, NVIDIA Container Toolkit, GPU architecture, `git`, Hugging Face CLI, token presence, and disk space. No network calls. | `run_script("scripts/validate_setup.py")` or `python scripts/validate_setup.py` |286| `scripts/.env.example` | Template for `HF_TOKEN` and optional `NGC_API_KEY`. **Keep the filled-in file OUTSIDE the checkout.** `Dockerfile.cosmos:29` is `COPY . /harmonizer-codebase`, and this repository ships no `.dockerignore`, so anything in the checkout — including a `.env` — is uploaded to the Docker daemon or remote builder and retained in build cache. A `.gitignore` does **not** scope build context and will not prevent this. | `install -m 600 scripts/.env.example ~/.harmonizer.env && set -a && . ~/.harmonizer.env && set +a` (the template is mode 0664 — a plain `cp` would leave your live tokens readable by other local users) |287288## References289290- [`references/inference.md`](references/inference.md) — container291 build, raw-base fallback, Blackwell patches, checkpoint292 download, `inference_pix2pix_turbo_harmonizer.py` flag matrix,293 non-temporal mode.294- [`references/evaluation.md`](references/evaluation.md) — paired295 `test_dataset/` layout and `evaluate_test_dataset.py` command, plus the296 upstream blockers that stop it running today.297- [`references/training.md`](references/training.md) — dataset298 download, training JSON format, multi-GPU `accelerate launch`299 recipe, fine-tuning flags, NuRec data-pair recipes.300- [`references/wrapper-image.md`](references/wrapper-image.md) —301 build and run the project image for repeat inference.302- [`references/troubleshooting.md`](references/troubleshooting.md)303 — extended diagnostic notes.304- [`references/teardown.md`](references/teardown.md) — cleanup305 inventory for images, code, Hugging Face caches, datasets,306 and outputs.307- Public code: <https://github.com/NVIDIA/harmonizer>308- Model card: <https://huggingface.co/nvidia/Harmonizer>309- Dataset: <https://huggingface.co/datasets/nvidia/Harmonizer-Dataset>310- Paper: <https://arxiv.org/abs/2602.24096>311- Project page: <https://research.nvidia.com/labs/sil/projects/diffusion-harmonizer/>312313## Limitations314315- **Rendered inputs only.** The model is tuned for316 neural-reconstruction renderings and object-insertion317 artifacts, not arbitrary real photos.318- **Primary model resolution is 576x1024.** The inference script319 maps resolution key `1024` to `1024x576`. Only `1024`, `960`,320 and `1360` are supported keys; `1024` matches the model-card321 operating point.322- **Temporal references and filename order matter.** The323 inference script enhances frames in plain lexically sorted order324 (zero-pad filenames to a fixed width) and325 feeds previous outputs back as temporal references. Use326 zero-padded frame numbers, or pass `--nontemporal` for327 unordered images.328- **Container builds can be large.** The runtime image, build329 cache, model weights, and optional dataset can exceed 100 GB.330- **Training is multi-GPU by default.** The README command331 assumes 8 GPUs with bf16 mixed precision.332- **Public code evolves.** The current README documents333 `inference_pix2pix_turbo_harmonizer.py` as the inference entry334 point; if a future checkout renames it, prefer the script and335 flags present in that checkout.336337## Troubleshooting (top 5)338339| Error / symptom | Most common cause |340|-----------------|-------------------|341| `docker: could not select device driver ... gpu` | NVIDIA Container Toolkit missing or Docker is not configured for the NVIDIA runtime. |342| `docker pull` 401 / 403 from `nvcr.io` | Docker is not authenticated to NGC, or the API key lacks container access. |343| `hf download ... 401 / 403` | Hitting a gated repo (`nvidia/Cosmos-Predict2-0.6B-Text2Image`, `nvidia/Harmonizer-Dataset`) without an accepted licence or a valid token. `nvidia/Harmonizer` itself is public and needs neither. |344| `diffusion_harmonizer.pkl` missing | Checkpoint download path is wrong or incomplete. Re-run `./download_checkpoints.sh` from the repo root. |345| Output files owned by `root` | The `docker run` omitted `-u $(id -u):$(id -g)`. |346347Full matrix in348[`references/troubleshooting.md`](references/troubleshooting.md).349350## Teardown351352A full workflow can leave large artifacts on disk: the Cosmos353image, project image, build cache, `harmonizer` code checkout,354Hugging Face model weights, optional dataset, evaluation355outputs, and enhanced frames. Reclaim them with the inventory356in [`references/teardown.md`](references/teardown.md). Do not357revoke `HF_TOKEN` or `NGC_API_KEY` as normal cleanup. Rotate a358token only if you suspect it was leaked.