workflow_wan_2_1_image_to_video_14b_720p
Imported workflow skill generated from Wan 2.1 Image to Video 14B 720p.json.
Family
video_t2v_i2v_avatar
Inputs
- Optional runtime overrides supported by
run(...):promptnegative_promptwidth,heightseed,steps,cfgsampler_name,scheduler,denoiseserver,headers,api_prefix
Outputs
- Returns JSON with:
statusprompt_idoutput_images
Model Requirements
clip:umt5_xxl_fp8_e4m3fn_scaled.safetensors->models/clipvae:wan_2.1_vae.safetensors->models/vaediffusion_model:wan2.1-i2v-14b-720p-Q4_0.gguf->models/diffusion_models
Custom Node Requirements
comfyui-videohelpersuite
Links
- https://discord.com/invite/gggpkVgBf3
- https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/clip_vision/clip_vision_h.safetensors?download=true
- https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors?download=true
- https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors?download=true
- https://huggingface.co/city96/Wan2.1-I2V-14B-720P-gguf/resolve/main/wan2.1-i2v-14b-720p-Q4_0.gguf?download=true
- https://huggingface.co/city96/Wan2.1-I2V-14B-720P-gguf/tree/main
- https://www.youtube.com/@pixaroma
Routing Metadata
- Family:
video_t2v_i2v_avatar - Input modalities:
image, text_prompt - Output modalities:
video/mp4 - Model families:
wan - Node count:
14 - Complexity score:
8 - Resource profile:
high - Estimated runtime:
slow (often 2-6 min depending on model/server load) - Max latent resolution hint:
NonexNone - Max sampler steps hint:
30
Detected Models
clip:umt5_xxl_fp8_e4m3fn_scaled.safetensors->models/clipvae:wan_2.1_vae.safetensors->models/vaediffusion_model:wan2.1-i2v-14b-720p-Q4_0.gguf->models/diffusion_models
Detected Custom Nodes
comfyui-videohelpersuite
Runtime Warnings
- Large model(s) detected; ensure enough VRAM and disk space.
- Uses custom nodes; missing nodes can cause validation/runtime failures.
- Video workflow: usually slower and VRAM-intensive than still-image workflows.