Datasheet for local AI50 models99 GPUsData read 2026-10-01
No adsNo tracking
Check my GPU →

Out of memory in ComfyUI: fixes in the order to try them

Fixes for CUDA out-of-memory errors in ComfyUI from the easiest to the last resort, plus why VRAM stays full after a run and what to do when it gets slow instead.

G.13FixesUpdated 2026-10-10

"CUDA out of memory" (or "HIP out of memory" on AMD) means the model plus its working memory did not fit on your GPU. Try these in order; most people are sorted by step 2 or 3.

1. Close what else is using the GPU

Browsers with hardware acceleration, games left running in the tray, a second ComfyUI window, video players. On Windows, Task Manager → Performance → GPU shows "Dedicated GPU memory" in use before you even start.

2. Use a smaller file of the same model

This is the big one. If you were on the 16-bit file, switch to the FP8 or INT8 file, or a Q8_0 GGUF. If 8-bit does not fit either, take a Q5 or Q4 GGUF and load it with Unet Loader (GGUF) from the ComfyUI-GGUF node pack. Which quant to pick →

3. Shrink the text encoder too

Text encoders like T5-XXL, UMT5 or Qwen-VL are several GB each. Use the FP8 or a GGUF version, or let the encoder run on the CPU. It only runs once per prompt, so on the CPU a T5-class encoder costs seconds; the 24–32B encoders of FLUX.2 [dev] and MiniMax H3 can take a minute or more.

4. Decode with the tiled VAE

If the error comes at the very end, the VAE decode is the problem, especially for video and big images. Replace VAE Decode with VAE Decode (Tiled).

5. Lower the resolution or the number of frames

Working memory grows with the number of pixels, and for video with the number of frames. Going from 1280×720 to 832×480, or from 121 frames to 81, frees a lot. Why video needs so much →

6. Update ComfyUI, and leave room for other apps

Since March 2026 ComfyUI's Dynamic VRAM is on by default for NVIDIA: it streams whatever does not fit from system RAM on its own, so an up-to-date install fixes many out-of-memory errors by itself. If other programs need GPU memory, start ComfyUI with --reserve-vram 1.5 to keep 1.5 GB free. The old --lowvram option does nothing when Dynamic VRAM is on; without it, it only moves the text encoder to the CPU.

7. Check your system RAM and page file

When VRAM runs out, the overflow goes to system RAM. If that is full too, older ComfyUI versions and non-NVIDIA setups fall back to the Windows page file and everything crawls, or crashes. (With Dynamic VRAM on NVIDIA, ComfyUI maps model files from disk instead and leans on the page file much less.) Big models want 32 GB of RAM; Wan 2.2 14B, FLUX.2 [dev] and MiniMax H3 want 64 GB or more. How much RAM →

Real numbers from my 32 GB PC: Wan 2.2 14B as GGUF Q8_0 pushed system RAM to 32.9 GB and Ming-Image to 31.3 GB — right at the limit. Every model timed on a 16 GB card →

VRAM stays full after a run — is that a memory leak?

Usually not. ComfyUI keeps the last models in VRAM on purpose, so the next run starts without loading them again. Task Manager then shows the card almost full even when nothing is running. That memory is handed back when another model needs it.

If you want it freed anyway:

  • Use Unload Models (or Unload Models and Execution Cache) from ComfyUI's command menu; ComfyUI-Manager has the same thing as the Free model and node cache button.
  • Start ComfyUI with --disable-smart-memory: models are moved out of VRAM after each run. Each run then starts a bit slower.
  • If system RAM keeps growing from run to run, try --cache-none. ComfyUI then keeps no node results between runs, which saves RAM but re-runs every node each time.

A real leak looks different: memory keeps climbing with every run of the same workflow until it crashes. Then the cause is usually a custom node. Update your custom nodes, or start ComfyUI without them to check.

Slow instead of out of memory?

On Windows, NVIDIA drivers can spill VRAM into “shared GPU memory” (normal RAM) instead of raising an error. The run then does not crash, it just becomes many times slower. Task Manager shows it under GPU → Shared GPU memory. If you would rather get a clear error and pick a smaller file, open NVIDIA Control Panel → Manage 3D settings → CUDA – Sysmem Fallback Policy and set it to Prefer No Sysmem Fallback.

Not sure which file fits your card in the first place? Check your GPU — it shows the largest file that fits entirely, for every model.