InternVideo2-CLIP (TAO video_clip)
Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
TAO task video_clip wraps OpenGVLab InternVideo2-CLIP L14. The PyTorch image provides train, evaluate, inference, export, and default_specs. TAO Deploy provides gen_trt_engine, TensorRT evaluate, and TensorRT inference.
Container images and per-action commands are in references/skill_info.yaml and references/tao-deploy-video-clip.skill_info.yaml. Starting specs are in references/spec_template_*.yaml.
Release note: The pinned PyTorch image is the TAO 7.2 FC build validated for Video-CLIP. The pinned tag resolves to OCI digest
sha256:faeb58559e1d87afd16453580999c178feb345f9dc87e35b8f40098e9604dd09; it includes PyAV 17.1.0 as the primary decoder and ONNXScript 0.7.1 for export, with decord absent. The TAO Deploy image is pinned independently becausegen_trt_engineand TensorRT-backed actions do not run in the PyTorch image.Known-broken images: interim builds cut before tao-pytorch commit
0cc31de4ship avideo_clippackage with nomodel.backbonessubmodule, sotrain/evaluate/inferencedie at import whilevideo_clip --helpstill exits 0. Images without PyAV also fail at data loading. Run both import checks in the preflight below before pulling data or launching a run.
Train Action Policy
AutoML is not packaged for this model skill. Always use direct video_clip actions even when a higher-level request mentions AutoML. Non-train actions stay in this skill.
Quick Start (local Docker)
Use the pinned TAO container declared in references/skill_info.yaml. Pull with NGC_KEY when the image is not cached locally.
VIDEO_CLIP_IMAGE_DEFAULT="nvcr.io/nvstaging/tao/tao-toolkit-pyt:v7.0.1-pyt2.1.0-py3-06" # versions-key: images.tao_toolkit.video_clip
VIDEO_CLIP_IMAGE="${VIDEO_CLIP_IMAGE:-$VIDEO_CLIP_IMAGE_DEFAULT}"
docker pull "$VIDEO_CLIP_IMAGE"
Expected workspace layout (host paths bind-mounted into the container):
workspace/
├── data/
│ ├── train.json # vadr1_chunks metadata; video_path = /data/videos/<name>.mp4
│ ├── val.json
│ ├── prompts.txt # one text prompt per line (inference)
│ └── videos/ # mp4 clips referenced by train.json / val.json
├── model/
│ ├── mobileclip_blt.pt # MobileCLIP weights for model.text_encoder
│ └── hf/ # offline InternVideo2_distillation_models snapshot (HF_HUB_OFFLINE=1)
├── specs/
│ ├── train.yaml
│ ├── evaluate.yaml
│ ├── inference.yaml
│ └── export.yaml
├── deploy_specs/
│ ├── gen_trt_engine.yaml
│ ├── evaluate.yaml
│ └── inference.yaml
└── results/
Docker options for all actions (skill-eval CI uses the same $WORKSPACE_DIR bind-mount pattern):
VIDEO_CLIP_IMAGE_DEFAULT="nvcr.io/nvstaging/tao/tao-toolkit-pyt:v7.0.1-pyt2.1.0-py3-06" # versions-key: images.tao_toolkit.video_clip
VIDEO_CLIP_IMAGE="${VIDEO_CLIP_IMAGE:-$VIDEO_CLIP_IMAGE_DEFAULT}"
RUN_ROOT="${RUN_ROOT:-$PWD}"
DOCKER_COMMON=(
--rm --gpus all --ipc=host --network=host
--shm-size=64g
--ulimit memlock=-1
--ulimit stack=67108864
-e WANDB_DISABLED=true
-e WANDB_MODE=disabled
-e HF_HUB_OFFLINE=1
-e HUGGINGFACE_HUB_CACHE=/model/hf
-e TRANSFORMERS_OFFLINE=1
-e TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1
-v "$RUN_ROOT/data:/data:ro"
-v "$RUN_ROOT/model:/model:ro"
-v "$RUN_ROOT/specs:/specs:ro"
-v "$RUN_ROOT/results:/results"
)
Preflight (host):
[ -f "$RUN_ROOT/model/mobileclip_blt.pt" ] || echo "MISSING: MobileCLIP weights"
[ -f "$RUN_ROOT/data/train.json" ] || echo "MISSING: train metadata"
[ -d "$RUN_ROOT/model/hf" ] || echo "MISSING: offline HF snapshot under model/hf"
docker run --rm "$VIDEO_CLIP_IMAGE" video_clip --help >/dev/null || echo "MISSING: video_clip in container"
docker run --rm "$VIDEO_CLIP_IMAGE" \
python -c "import nvidia_tao_pytorch.multimodal.video_clip.model.adapters.internvideo2clip" \
>/dev/null 2>&1 || echo "BROKEN IMAGE: video_clip package is incomplete (missing model.backbones) — stop, see Release note"
docker run --rm "$VIDEO_CLIP_IMAGE" python -c "import av; print(av.__version__)" \
>/dev/null 2>&1 || echo "BROKEN IMAGE: PyAV is missing — use the pinned FC image; do not install decord"
nvidia-smi >/dev/null 2>&1 || echo "note: no GPU visible"
Train:
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
video_clip train -e /specs/train.yaml results_dir=/results
Evaluate:
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
video_clip evaluate -e /specs/evaluate.yaml results_dir=/results
Inference:
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
video_clip inference -e /specs/inference.yaml results_dir=/results
Export:
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
video_clip export -e /specs/export.yaml results_dir=/results
TensorRT Deploy
Use the independently pinned TAO Deploy image after PyTorch export. Read references/tao-deploy-video-clip.md before running the deploy actions; its templates cover the engine build, retrieval evaluation, and embedding inference contracts.
VIDEO_CLIP_DEPLOY_IMAGE_DEFAULT="nvcr.io/nvstaging/tao/tao-toolkit-deploy:7.2.0-rc-47-multiarch" # versions-key: images.tao_toolkit.video_clip_deploy
VIDEO_CLIP_DEPLOY_IMAGE="${VIDEO_CLIP_DEPLOY_IMAGE:-$VIDEO_CLIP_DEPLOY_IMAGE_DEFAULT}"
# Verify the independently pinned image before staging artifacts or using a GPU.
docker run --rm "$VIDEO_CLIP_DEPLOY_IMAGE" video_clip gen_trt_engine --help >/dev/null || \
{ echo "BROKEN IMAGE: Video-CLIP deploy entrypoint is unavailable" >&2; exit 1; }
docker run --gpus all --rm --shm-size=16g \
-v "$RUN_ROOT/deploy_specs:/specs:ro" \
-v "$RUN_ROOT/results/export:/models:ro" \
-v "$RUN_ROOT/data:/data:ro" \
-v "$RUN_ROOT/results/deploy:/results" \
"$VIDEO_CLIP_DEPLOY_IMAGE" \
video_clip gen_trt_engine -e /specs/gen_trt_engine.yaml
Keep the exported ONNX file, its matching *_config.yaml, and its matching *_tokenizer/ directory together. gen_trt_engine copies the sidecars beside the engine so TensorRT evaluate and inference can reconstruct preprocessing and tokenization.
Quick Start (virtualenv — local dev hosts)
On hosts with a tao-pytorch checkout and tao-cli venv (for example rtdetr-pytorch), run through tao-run-on-virtualenv instead of Docker:
export VENV="${VENV:?set to your tao-cli virtualenv (must contain bin/video_clip)}"
export TAO_PYTORCH_ROOT="${TAO_PYTORCH_ROOT:?set to your tao-pytorch checkout with multimodal/video_clip}"
export PATH="$VENV/bin:$PATH"
export PYTHONPATH="$TAO_PYTORCH_ROOT:$TAO_PYTORCH_ROOT/tao-core:$PYTHONPATH"
export HF_HOME="${HF_HOME:-$PWD/hf_cache}"
export HF_HUB_OFFLINE="${HF_HUB_OFFLINE:-1}"
export WANDB_DISABLED=true
export WANDB_MODE=disabled
Copy a template spec from references/spec_template_*.yaml, fill checkpoint paths and metadata, then run:
video_clip train -e /path/to/train.yaml
video_clip evaluate -e /path/to/evaluate.yaml
video_clip inference -e /path/to/inference.yaml
video_clip export -e /path/to/export.yaml
Credentials
- NGC_KEY — pull the pinned TAO container from
nvcr.iowhen it is not cached locally. - HF_TOKEN (online Hugging Face download only): Hugging Face read token used
when
model.vision_encoderandmodel.clip_headarenulland the resolver downloads the InternVideo2 snapshot named bymodel.internvideo2clip_hf_id. For offline CI/eval, stage the complete S3/local snapshot and point both fields at the local files; no Hugging Face token is then required.
Treat tokens as secrets. Export them into the environment or pass them through an
--env-file of bare KEY=value lines, rather than inlining values into generated
spec YAML, command lines, or anything written under results_dir.
Data format (vadr1_chunks)
Metadata is a top-level JSON list of video records. Each record has video_path, split, and nested chunks[] with caption fields (queries, action_queries, anomaly_queries, dense_caption, scene_caption).
Point dataset.*.video_text.metadata at the user's JSON files (for example /data/train.json and /data/val.json in the templates). Each video_path must be an absolute path resolvable inside the runtime — for Docker, remap host clips to container paths like /data/videos/<name>.mp4 and set data_root: null unless using path_prefix_mapping.
Smoke overrides
For a short functional check (for example 2 epochs, 1 GPU, small batch):
train.num_epochs: 2
train.num_gpus: 1
train.gpu_ids: [0]
train.optim.warmup_steps: 10
dataset.train.batch_size: 2
dataset.val.batch_size: 2
dataset.train.num_workers: 4
dataset.val.num_workers: 4
dataset.train.video_text.caption_fields: [queries, action_queries, anomaly_queries]
dataset.train.video_text.caption_mode: first
dataset.metrics.mode: classification
Use dataset.metrics.mode: retrieval only when dataset.val.video_text.relevance_file is provided.
Inference
inference.mode: embeddingswritesvideo_embeddings.h5andtext_embeddings.h5underresults_dir.- Provide
inference.query.text_file(one prompt per line) and/orinference.query.input_texts. - Gallery videos come from
dataset.inference.video_text.metadata.
Export
export.encoder_type: combined produces the image-and-text ONNX consumed by the Video-CLIP deploy workflow. Keep export.batch_size: -1 for symbolic/dynamic batch dimensions; a positive value produces a fixed-batch ONNX. Export also writes matching *_config.yaml and *_tokenizer/ sidecars; preserve all three artifacts. Default opset is 23 on the vendor branch. Export requires a trained .pth at export.checkpoint.
LoRA
For vision-LoRA runs, start from tao-pytorch experiment_spec_lora.yaml or add a top-level peft: block (see shipped spec comments). Merge LoRA before export when checkpoints contain lora_* keys.
Common pitfalls
PATHmust prefer$VENV/binon virtualenv hosts so child processes resolve the venv Python.evaluateusesdataset.val, not a separate test split — Lightning stage"test"still loads val metadata.- Eval precision: match
train.precision(typicallybf16) or flash-attn paths may fail under fp32 eval. - Empty
action_querieson normal chunks become literal"Normal"positives during training; excludeNormal/Abnormalindataset.metrics.exclude_categoriesfor classification eval. - No hard-negative / explicit-neg training on the vendor branch unless the spec and branch explicitly enable it.
video_clip --helpis not a health check. It exits 0 on an image whosevideo_clippackage is missingmodel.backbones; only the import smoke check in the preflight catches it.- PyAV is the primary Video-CLIP decoder in TAO 7.2. The loader can fall back to the image's FFmpeg CLI and OpenCV support. A missing
avimport is an image defect: use the pinned FC image. Do not add or force-install decord, because it is not part of the supported TAO 7.2 decode contract. - TensorRT actions use TAO Deploy.
gen_trt_engine, TensorRTevaluate, and TensorRTinferencemust use the independently pinned deploy image and deploy templates, not the PyTorch image/specs. - Deploy sidecars are required. TensorRT evaluation and text inference need the exported
*_config.yamland*_tokenizer/beside the engine. Keep them with the ONNX input so engine generation can copy them automatically. - PyTorch ≥ 2.6 defaults to
torch.load(weights_only=True)and rejects the TAO checkpoint’s numpy dtype objects with_pickle.UnpicklingError.TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1is set inDOCKER_COMMONabove; keep it forevaluate,inference, andexport. model/hf/is a snapshot, not an HF hub cache. Offline packs stage InternVideo2 weights at repo-relative paths (stage1/L14/L14_dist_1B_stage2/pytorch_model.bin,clip/L14/pytorch_model.bin), whileHUGGINGFACE_HUB_CACHEexpects amodels--<org>--<repo>/snapshots/<sha>/tree. Leavingmodel.vision_encoder/model.clip_headatnullsends asset resolution tohf_hub_downloadand fails underHF_HUB_OFFLINE=1— point both at the files directly.- CI / skill-eval: stage weights + remapped JSON under
$WORKSPACE_DIRfrom S3; do not rely on HuggingFace LFS downloads at eval time.
References
- Shipped defaults:
tao-pytorchnvidia_tao_pytorch/multimodal/video_clip/experiment_specs/(when developing from source) - TAO Deploy workflow:
references/tao-deploy-video-clip.md