ComfyUI launch/performance flags
Overview
CLI flags passed to main.py control ComfyUI's runtime behavior
(e.g. python main.py --reserve-vram 2 --use-sage-attention). The three that
matter most for making a graph run rather than OOM or crawl are the
VRAM strategy, the attention backend, and the cache mode. This skill
is the decision matrix for choosing them.
⚠️ Verification note (August 2026). Every flag below was checked against
upstream comfy/cli_args.py
on current master. ComfyUI adds/renames flags often — when in doubt run
python main.py --help in the target install and prefer that over this list.
--enable-triton-backend / --disable-triton-backend ARE ComfyUI main.py
flags on master (they used to be documented as SwarmUI-only; that is stale).
--use-ck-attention is kitchen INT8 attention — no sageattention wheel.
A June ComfyUI checkout still pins comfy-kitchen 0.2.10 and lacks
--use-ck-attention; kitchen action:"status" reports ComfyUI-side flag
support, not only the kitchen version. Use kitchen / panel_kitchen to see
what this GPU can actually run.
How to apply today. The MCP's restart_comfyui (with action: "start")
currently replays the exact argv of the previous run. It does not compose
fresh flags. So set these when you launch ComfyUI yourself (the
python main.py … line, a run.bat/shell alias, or the SwarmUI backend args
box), and the tool will preserve them on restart. Injecting flags through the
tool is a tracked follow-up.
Decide first: which flag do you need?
Symptom ▶ Flag(s) to try
─────────────────────────────────────────────────────────────────────────────
CUDA out of memory, long video (LTX 2 / WAN) ▶ --novram (+ --cache-none)
OOM, still want models resident when they fit ▶ --reserve-vram N then --disable-smart-memory
GPU slows to a crawl, spills into "shared GPU ▶ --reserve-vram 2..4
memory" (Windows WDDM) mid-run
RAM blows up switching between models, or a huge ▶ --cache-none
text encoder (FLUX 2 / Mistral) won't unload
Plenty of VRAM (48GB+), want max throughput ▶ --gpu-only or --highvram
Want faster sampling on NVIDIA ▶ --use-ck-attention if kitchen INT8 is available (skip the sage wheel); else --use-sage-attention
Z-Image produces BLACK / wrong output ▶ --use-pytorch-cross-attention (NOT sage)
Sage gives black output on some models ▶ --use-pytorch-cross-attention (or fix dtype)
ROCm, kitchen present, triton ≥ 3.7 ▶ --enable-triton-backend
VRAM strategy and attention backend are each mutually exclusive groups, so
pass at most one from each. You can combine one VRAM flag + one attention flag +
one cache flag (e.g. --novram --use-sage-attention --cache-none).
VRAM strategy (mutually exclusive)
| Flag |
What it does |
Use when |
--gpu-only |
Keep everything (incl. text encoders) on GPU |
48GB+ card, single model, max speed |
--highvram |
Keep models resident in VRAM after use |
High-VRAM card, repeated runs of one model |
| (default) |
ComfyUI's smart offload |
Most setups — try this first |
--lowvram |
Offload text encoders / parts to CPU |
Mid card OOMing on load |
--novram |
Extreme offload — minimal VRAM footprint |
OOM on long video / huge models; pair with --cache-none |
--cpu |
Everything on CPU (very slow) |
No usable CUDA GPU only |
Modifiers (combine with the above):
--reserve-vram N reserves N GB for the OS and other apps. It is the fix for the
Windows failure mode where the GPU quietly starts using shared VRAM and
throughput collapses. Typical 2 to 4; bump to 10 for heavy video decode.
--disable-smart-memory forces aggressive offload to regular RAM instead
of keeping models cached in VRAM. Reach for this when a run gets stuck or
OOMs intermittently. Slightly slower, much more reliable.
--async-offload enables async weight offload streams (default on where
supported); --disable-async-offload turns it off if it misbehaves.
Attention backend (mutually exclusive)
| Flag |
Notes |
--use-ck-attention |
Comfy Kitchen INT8 attention. No sageattention wheel. Needs comfy-kitchen present and int8_attention_is_available() on this GPU. Prefer this over the sage wheel-matching install when kitchen action:"status" says INT8 is available. Restart required. |
--use-sage-attention |
Quantized SageAttention kernel, ~20–40% faster sampling. Needs the sageattention package installed and version-matched — see triton-sageattention. Skip this dance when --use-ck-attention is available. |
--use-flash-attention |
FlashAttention kernels. Needs flash-attn built for your torch/CUDA. |
--enable-triton-backend / --disable-triton-backend |
Enable or disable the comfy-kitchen triton backend. ComfyUI master flags (not SwarmUI-only). ROCm hosts with kitchen + triton ≥ 3.7 want --enable-triton-backend. Restart required. |
--use-pytorch-cross-attention |
PyTorch SDPA. Highest quality, always available, no extra deps. The safe default and the correct fallback. |
--use-split-cross-attention / --use-quad-cross-attention |
Memory-optimized math attention for older/low-VRAM cards. |
Two gotchas worth memorizing:
- Z-Image + Sage = broken. Z-Image (Turbo/Base) does not sample
correctly under
--use-sage-attention; you get black or garbled output.
Launch Z-Image with --use-pytorch-cross-attention instead. See
z-image-txt2img.
- Sage black output on other models. If a model outputs black only with
Sage, either switch to
--use-pytorch-cross-attention, or (SwarmUI) set
Advanced Sampling → Preferred DType = Default (16-bit). Sage-on vs Sage-off
also produces slightly different images, so expect non-identical seeds.
When a graph hard-crashes with No module named 'sageattention' /
triton: unavailable, the fix is the sdpa / no-compile fallback in
triton-sageattention, not this flag.
Cache mode (mutually exclusive)
| Flag |
Effect |
(default --cache-ram) |
Cache results under RAM pressure |
--cache-classic |
Aggressive result caching |
--cache-lru N |
Keep at most N node results (LRU) |
--cache-none |
Cache nothing — re-executes every node; lowest RAM/VRAM. Essential when switching between dual models or when a giant text encoder (FLUX 2's Mistral) must fully unload. |
Speed / precision
--fast enables experimental, potentially quality-degrading
optimizations. Accepts specific PerformanceFeature values:
fp16_accumulation, fp8_matrix_mult, cublas_ops, autotune. Bare --fast
turns them all on. Test output quality before committing to it.
- UNet/VAE/text-encoder dtype casts exist too
(
--fp8_e4m3fn-unet, --fp16-unet, --bf16-unet, --fp32-unet, …) for
forcing a compute precision. Usually the model or loader picks the right one, so
only reach for these to work around a specific dtype error.
Recommended combos (recipes)
Long video OOM (LTX 2 / WAN, 24GB): --novram --cache-none
(add --disable-smart-memory if it stalls)
Windows shared-VRAM creep: --reserve-vram 3
FLUX 2 / huge text-encoder swaps: --cache-none
High-VRAM throughput (48GB+): --gpu-only (or --highvram)
Fast NVIDIA sampling (most models): --use-ck-attention (if kitchen INT8 is available)
--use-sage-attention (otherwise; needs the wheel)
Z-Image (any): --use-pytorch-cross-attention
ROCm + kitchen + triton ≥ 3.7: --enable-triton-backend
Cross-refs: video OOM specifics in
ltxv2-video / wan-t2v-video;
per-model VRAM math in troubleshooting and
model-compatibility.
Acceleration stack & GPU coverage (context)
The attention/compile accelerators are version-locked to your exact
torch + CUDA + Python. A mismatched wheel doesn't just fail to import; it can
break the torch install. A known-good, mutually-compatible stack for late-2025 /
2026 NVIDIA (including Blackwell / RTX 5000, sm_120) looks like:
| Component |
Role |
Notes |
| Torch + CUDA |
base |
e.g. Torch 2.9.x on CUDA 12.8/13; use the wheel index matching your driver |
| Triton |
torch.compile / inductor |
Windows: triton-windows (woct0rdho) |
| SageAttention |
--use-sage-attention |
wheel matched to torch/CUDA/python |
| FlashAttention |
--use-flash-attention |
built per torch/CUDA/python |
| xFormers |
memory-efficient attention |
optional |
| InsightFace |
FaceID / IP-Adapter / ReActor |
onnxruntime-gpu alongside |
Operational facts worth carrying:
- No system-wide CUDA toolkit is required to run ComfyUI. An up-to-date
NVIDIA driver plus prebuilt wheels is enough. A full CUDA/MSVC/cuDNN toolchain is
only needed to compile kernels yourself.
- For broad arch coverage when building wheels,
TORCH_CUDA_ARCH_LIST=7.5;8.0;8.6;8.9;9.0;10.0;12.0+PTX spans RTX 20xx→50xx
and datacenter (A100/H100/B200). +PTX lets newer archs JIT.
- DeepSpeed has no wheels for Python 3.13, and several accel wheels lag the
newest Python. 3.10 to 3.12 is the safe range for the full stack.
- Clear the Triton cache (
~/.triton / %USERPROFILE%\.triton and temp)
when you hit stale-kernel Triton errors after an upgrade.
- Prefer
uv pip install over pip for the venv. Resolves and downloads are
dramatically faster. install_comfyui already supports this via preferUv.
- A single bad custom node can crash all of ComfyUI at startup. Install and test
acceleration and new node packs on a fresh/known-good install, not before a
deadline. See
troubleshooting.
Quantization quick take
- FP8-scaled (per-tensor scaled) is markedly higher quality than plain
base FP8, ~half the size of BF16, and usually faster.
- Prefer FP8-scaled over GGUF when you have enough system RAM. ComfyUI's
block-swap streams from RAM, so BF16/FP8 can run on 24GB GPUs given ample RAM.
Fall back to GGUF (Q8→Q4) only when RAM is the constraint.
- NVFP4 / NVFP8 are markedly faster on Blackwell (RTX 5000) at near-BF16
quality for supported models; LoRA support on NVFP4 is still partial.
Sources
- Official: ComfyUI CLI args at https://github.com/comfyanonymous/ComfyUI/blob/master/comfy/cli_args.py (
--use-ck-attention, --enable-triton-backend, --disable-triton-backend, --fast); hardware gates in comfy/model_management.py (supports_fp8_compute SM ≥ 8.9, supports_nvfp4_compute / supports_mxfp8_compute SM ≥ 10.0); kitchen backends in the comfy-kitchen README https://github.com/Comfy-Org/comfy-kitchen
- Empirical: operational flag/stack recipes distilled from community auto-installer changelogs (SECourses); flags cross-checked against upstream above. The SwarmUI-only note for
--enable-triton-backend is retracted as of ComfyUI master.
1---2name: comfyui-launch-flags3description: Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed. The full decision matrix for OOM (--novram / --cache-none / --disable-smart-memory), shared-VRAM creep on Windows (--reserve-vram N), model-switching with big text encoders (--cache-none), high-VRAM throughput (--gpu-only / --highvram), and attention-backend selection (--use-sage-attention for speed, --use-pytorch-cross-attention as the highest-quality / Z-Image-safe fallback). Also the acceleration-stack + Blackwell/RTX 5000 (sm_120) notes. Use when a graph OOMs (especially long video like LTX 2 / WAN), when the GPU spills into shared VRAM and slows to a crawl, when switching between models eats all RAM, when Z-Image produces black/garbled output under Sage, or when deciding which attention backend to launch with. Flag names verified against upstream comfy/cli_args.py; see Sources.4---56# ComfyUI launch/performance flags78## Overview910CLI flags passed to `main.py` control ComfyUI's runtime behavior11(e.g. `python main.py --reserve-vram 2 --use-sage-attention`). The three that12matter most for making a graph *run* rather than OOM or crawl are the13VRAM strategy, the attention backend, and the cache mode. This skill14is the decision matrix for choosing them.1516> ⚠️ **Verification note (August 2026).** Every flag below was checked against17> upstream [`comfy/cli_args.py`](https://github.com/comfyanonymous/ComfyUI/blob/master/comfy/cli_args.py)18> on current master. ComfyUI adds/renames flags often — when in doubt run19> `python main.py --help` in the target install and prefer that over this list.20> **`--enable-triton-backend` / `--disable-triton-backend` ARE ComfyUI `main.py`21> flags on master** (they used to be documented as SwarmUI-only; that is stale).22> `--use-ck-attention` is kitchen INT8 attention — no `sageattention` wheel.23> A June ComfyUI checkout still pins comfy-kitchen 0.2.10 and lacks24> `--use-ck-attention`; `kitchen` action:"status" reports ComfyUI-side flag25> support, not only the kitchen version. Use `kitchen` / `panel_kitchen` to see26> what this GPU can actually run.2728> How to apply today. The MCP's `restart_comfyui` (with `action: "start"`)29> currently *replays the exact argv of the previous run*. It does not compose30> fresh flags. So set these when you launch ComfyUI yourself (the31> `python main.py …` line, a `run.bat`/shell alias, or the SwarmUI backend args32> box), and the tool will preserve them on restart. Injecting flags through the33> tool is a tracked follow-up.3435---3637## Decide first: which flag do you need?3839```40Symptom ▶ Flag(s) to try41─────────────────────────────────────────────────────────────────────────────42CUDA out of memory, long video (LTX 2 / WAN) ▶ --novram (+ --cache-none)43OOM, still want models resident when they fit ▶ --reserve-vram N then --disable-smart-memory44GPU slows to a crawl, spills into "shared GPU ▶ --reserve-vram 2..445 memory" (Windows WDDM) mid-run46RAM blows up switching between models, or a huge ▶ --cache-none47 text encoder (FLUX 2 / Mistral) won't unload48Plenty of VRAM (48GB+), want max throughput ▶ --gpu-only or --highvram49Want faster sampling on NVIDIA ▶ --use-ck-attention if kitchen INT8 is available (skip the sage wheel); else --use-sage-attention50Z-Image produces BLACK / wrong output ▶ --use-pytorch-cross-attention (NOT sage)51Sage gives black output on some models ▶ --use-pytorch-cross-attention (or fix dtype)52ROCm, kitchen present, triton ≥ 3.7 ▶ --enable-triton-backend53```5455VRAM strategy and attention backend are each mutually exclusive groups, so56pass at most one from each. You can combine one VRAM flag + one attention flag +57one cache flag (e.g. `--novram --use-sage-attention --cache-none`).5859---6061## VRAM strategy (mutually exclusive)6263| Flag | What it does | Use when |64|------|--------------|----------|65| `--gpu-only` | Keep everything (incl. text encoders) on GPU | 48GB+ card, single model, max speed |66| `--highvram` | Keep models resident in VRAM after use | High-VRAM card, repeated runs of one model |67| *(default)* | ComfyUI's smart offload | Most setups — try this first |68| `--lowvram` | Offload text encoders / parts to CPU | Mid card OOMing on load |69| `--novram` | Extreme offload — minimal VRAM footprint | OOM on long video / huge models; pair with `--cache-none` |70| `--cpu` | Everything on CPU (very slow) | No usable CUDA GPU only |7172Modifiers (combine with the above):7374- `--reserve-vram N` reserves N GB for the OS and other apps. It is the fix for the75 Windows failure mode where the GPU quietly starts using shared VRAM and76 throughput collapses. Typical `2` to `4`; bump to `10` for heavy video decode.77- `--disable-smart-memory` forces aggressive offload to regular RAM instead78 of keeping models cached in VRAM. Reach for this when a run gets *stuck* or79 OOMs intermittently. Slightly slower, much more reliable.80- `--async-offload` enables async weight offload streams (default on where81 supported); `--disable-async-offload` turns it off if it misbehaves.8283---8485## Attention backend (mutually exclusive)8687| Flag | Notes |88|------|-------|89| `--use-ck-attention` | Comfy Kitchen INT8 attention. **No `sageattention` wheel.** Needs comfy-kitchen present and `int8_attention_is_available()` on this GPU. Prefer this over the sage wheel-matching install when `kitchen` action:"status" says INT8 is available. Restart required. |90| `--use-sage-attention` | Quantized SageAttention kernel, ~20–40% faster sampling. Needs the `sageattention` package installed and version-matched — see [`triton-sageattention`](../triton-sageattention/SKILL.md). Skip this dance when `--use-ck-attention` is available. |91| `--use-flash-attention` | FlashAttention kernels. Needs `flash-attn` built for your torch/CUDA. |92| `--enable-triton-backend` / `--disable-triton-backend` | Enable or disable the comfy-kitchen **triton** backend. ComfyUI master flags (not SwarmUI-only). ROCm hosts with kitchen + triton ≥ 3.7 want `--enable-triton-backend`. Restart required. |93| `--use-pytorch-cross-attention` | PyTorch SDPA. **Highest quality, always available, no extra deps.** The safe default and the correct fallback. |94| `--use-split-cross-attention` / `--use-quad-cross-attention` | Memory-optimized math attention for older/low-VRAM cards. |9596Two gotchas worth memorizing:97981. **Z-Image + Sage = broken.** Z-Image (Turbo/Base) does not sample99 correctly under `--use-sage-attention`; you get black or garbled output.100 Launch Z-Image with `--use-pytorch-cross-attention` instead. See101 [`z-image-txt2img`](../z-image-txt2img/SKILL.md).1022. **Sage black output on other models.** If a model outputs black *only* with103 Sage, either switch to `--use-pytorch-cross-attention`, or (SwarmUI) set104 Advanced Sampling → Preferred DType = Default (16-bit). Sage-on vs Sage-off105 also produces *slightly different* images, so expect non-identical seeds.106107> When a graph hard-crashes with `No module named 'sageattention'` /108> `triton: unavailable`, the fix is the sdpa / no-compile fallback in109> [`triton-sageattention`](../triton-sageattention/SKILL.md), not this flag.110111---112113## Cache mode (mutually exclusive)114115| Flag | Effect |116|------|--------|117| *(default `--cache-ram`)* | Cache results under RAM pressure |118| `--cache-classic` | Aggressive result caching |119| `--cache-lru N` | Keep at most N node results (LRU) |120| `--cache-none` | Cache nothing — re-executes every node; **lowest RAM/VRAM**. Essential when switching between dual models or when a giant text encoder (FLUX 2's Mistral) must fully unload. |121122---123124## Speed / precision125126- `--fast` enables experimental, potentially quality-degrading127 optimizations. Accepts specific `PerformanceFeature` values:128 `fp16_accumulation`, `fp8_matrix_mult`, `cublas_ops`, `autotune`. Bare `--fast`129 turns them all on. Test output quality before committing to it.130- UNet/VAE/text-encoder dtype casts exist too131 (`--fp8_e4m3fn-unet`, `--fp16-unet`, `--bf16-unet`, `--fp32-unet`, …) for132 forcing a compute precision. Usually the model or loader picks the right one, so133 only reach for these to work around a specific dtype error.134135---136137## Recommended combos (recipes)138139```140Long video OOM (LTX 2 / WAN, 24GB): --novram --cache-none141 (add --disable-smart-memory if it stalls)142Windows shared-VRAM creep: --reserve-vram 3143FLUX 2 / huge text-encoder swaps: --cache-none144High-VRAM throughput (48GB+): --gpu-only (or --highvram)145Fast NVIDIA sampling (most models): --use-ck-attention (if kitchen INT8 is available)146 --use-sage-attention (otherwise; needs the wheel)147Z-Image (any): --use-pytorch-cross-attention148ROCm + kitchen + triton ≥ 3.7: --enable-triton-backend149```150151Cross-refs: video OOM specifics in152[`ltxv2-video`](../ltxv2-video/SKILL.md) / [`wan-t2v-video`](../wan-t2v-video/SKILL.md);153per-model VRAM math in [`troubleshooting`](../troubleshooting/SKILL.md) and154[`model-compatibility`](../model-compatibility/SKILL.md).155156---157158## Acceleration stack & GPU coverage (context)159160The attention/compile accelerators are version-locked to your exact161torch + CUDA + Python. A mismatched wheel doesn't just fail to import; it can162break the torch install. A known-good, mutually-compatible stack for late-2025 /1632026 NVIDIA (including Blackwell / RTX 5000, `sm_120`) looks like:164165| Component | Role | Notes |166|-----------|------|-------|167| Torch + CUDA | base | e.g. Torch 2.9.x on CUDA 12.8/13; use the wheel index matching your driver |168| Triton | `torch.compile` / inductor | Windows: `triton-windows` (woct0rdho) |169| SageAttention | `--use-sage-attention` | wheel matched to torch/CUDA/python |170| FlashAttention | `--use-flash-attention` | built per torch/CUDA/python |171| xFormers | memory-efficient attention | optional |172| InsightFace | FaceID / IP-Adapter / ReActor | `onnxruntime-gpu` alongside |173174Operational facts worth carrying:175176- No system-wide CUDA toolkit is required to *run* ComfyUI. An up-to-date177 NVIDIA driver plus prebuilt wheels is enough. A full CUDA/MSVC/cuDNN toolchain is178 only needed to *compile* kernels yourself.179- For broad arch coverage when building wheels,180 `TORCH_CUDA_ARCH_LIST=7.5;8.0;8.6;8.9;9.0;10.0;12.0+PTX` spans RTX 20xx→50xx181 and datacenter (A100/H100/B200). `+PTX` lets newer archs JIT.182- DeepSpeed has no wheels for Python 3.13, and several accel wheels lag the183 newest Python. 3.10 to 3.12 is the safe range for the full stack.184- Clear the Triton cache (`~/.triton` / `%USERPROFILE%\.triton` and temp)185 when you hit stale-kernel Triton errors after an upgrade.186- Prefer `uv pip install` over pip for the venv. Resolves and downloads are187 dramatically faster. `install_comfyui` already supports this via `preferUv`.188- A single bad custom node can crash all of ComfyUI at startup. Install and test189 acceleration and new node packs on a fresh/known-good install, not before a190 deadline. See [`troubleshooting`](../troubleshooting/SKILL.md).191192## Quantization quick take193194- FP8-*scaled* (per-tensor scaled) is markedly higher quality than plain195 base FP8, ~half the size of BF16, and usually faster.196- Prefer FP8-scaled over GGUF when you have enough system RAM. ComfyUI's197 block-swap streams from RAM, so BF16/FP8 can run on 24GB GPUs given ample RAM.198 Fall back to GGUF (Q8→Q4) only when RAM is the constraint.199- NVFP4 / NVFP8 are markedly faster on Blackwell (RTX 5000) at near-BF16200 quality for supported models; LoRA support on NVFP4 is still partial.201202---203204## Sources205206- **Official:** ComfyUI CLI args at https://github.com/comfyanonymous/ComfyUI/blob/master/comfy/cli_args.py (`--use-ck-attention`, `--enable-triton-backend`, `--disable-triton-backend`, `--fast`); hardware gates in `comfy/model_management.py` (`supports_fp8_compute` SM ≥ 8.9, `supports_nvfp4_compute` / `supports_mxfp8_compute` SM ≥ 10.0); kitchen backends in the comfy-kitchen README https://github.com/Comfy-Org/comfy-kitchen207- **Empirical:** operational flag/stack recipes distilled from community auto-installer changelogs (SECourses); flags cross-checked against upstream above. The SwarmUI-only note for `--enable-triton-backend` is retracted as of ComfyUI master.