curated_qwenvl_video_to_text
Curated workflow skill generated from QwenVL Video to Text.json.
Capability Family
video_t2v_i2v_avatar
Inputs
- Optional runtime overrides supported by
run(...):promptnegative_promptwidth,heightseed,steps,cfgsampler_name,scheduler,denoiseserver,headers,api_prefix
Outputs
- Returns JSON with:
statusprompt_idoutput_images(includes image/video entries reported by Comfy history)
Model Requirements
- None detected from loader nodes.
Custom Node Requirements
comfyui-qwenvlcomfyui-videohelpersuite
Links Extracted From Workflow Notes
- https://discord.com/invite/gggpkVgBf3
- https://github.com/1038lab/ComfyUI-QwenVL
- https://www.youtube.com/@pixaroma
Source
- Original:
comfy-data/workflows/QwenVL Video to Text.json
Routing Metadata
- Family:
video_t2v_i2v_avatar - Input modalities:
video - Output modalities:
application/json - Model families:
qwen, wan - Node count:
5 - Complexity score:
5 - Resource profile:
medium - Estimated runtime:
moderate (about 30-120s depending on server) - Max latent resolution hint:
NonexNone - Max sampler steps hint:
None
Detected Models
- None detected.
Detected Custom Nodes
comfyui-qwenvlcomfyui-videohelpersuite
Runtime Warnings
- Uses custom nodes; missing nodes can cause validation/runtime failures.
- Video workflow: usually slower and VRAM-intensive than still-image workflows.