Datasheet for local AI50 models99 GPUsData read 2026-10-01
No adsNo tracking
Check my GPU →

umt5_xxl_fp8_e4m3fn_scaled.safetensors

UMT5-XXL FP8 (scaled) — the text encoder for Wan 2.1 14B, Wan 2.1 1.3B, Wan 2.1 I2V 480P and 9 more models. 6.7 GB. In ComfyUI it goes in models/text_encoders.

6.7 GBmodels/text_encoders12 models
Fileumt5_xxl_fp8_e4m3fn_scaled.safetensors
What it isUMT5-XXL FP8 (scaled) (text encoder)
Size6.7 GB
FolderComfyUI/models/text_encoders/
Loader nodea text-encoder loader: Load CLIP (CLIPLoader) for one encoder, DualCLIPLoader when the model takes two (FLUX.1: CLIP-L + T5-XXL)
SourceComfy-Org/Wan_2.1_ComfyUI_repackaged on Hugging Face

Download umt5_xxl_fp8_e4m3fn_scaled.safetensors (6.7 GB)

Where to put it

Put the file in ComfyUI/models/text_encoders/ (in the portable version: ComfyUI_windows_portable/ComfyUI/models/text_encoders/). Older guides say models/clip: that folder still works. Then press R in ComfyUI, or restart it, so the file shows up in a text-encoder loader: Load CLIP (CLIPLoader) for one encoder, DualCLIPLoader when the model takes two (FLUX.1: CLIP-L + T5-XXL).

If a workflow says “Value not in list” for this node, the file name in the workflow is different from yours, or the file is in another folder. Click the node and pick the file from the list.

What it does

A text encoder turns your prompt into numbers the model understands. It runs once per prompt, then ComfyUI can move it out of VRAM to make room for the image model — so it costs load time and system RAM more than speed.

This is a smaller FP8 version. It loads faster and needs about half the RAM of the 16-bit file; the difference in the pictures is usually hard to see. It is the version this site lists by default.

FP8 files load on every NVIDIA card in ComfyUI. On a Mac (Apple Silicon) FP8 does not load — use the 16-bit file there.

Models that use umt5_xxl_fp8_e4m3fn_scaled.safetensors

ModelTypeFits fromRuns well from
SCAIL-2 (character animation)video13 GB22 GB
Wan 2.1 I2V 14B 480Pvideo13 GB21 GB
Wan 2.1 I2V 14B 720Pvideo16 GB24 GB
Wan 2.1 T2V 1.3Bvideo4 GB5 GB
Wan 2.1 T2V 14Bvideo12 GB19 GB
Wan 2.1 VACE 14Bvideo13 GB24 GB
Wan 2.2 Animate 14Bvideo12 GB23 GB
Wan 2.2 I2V A14Bvideo10 GB19 GB
Wan 2.2 S2V 14Bvideo15 GB22 GB
Wan 2.2 T2V A14Bvideo10 GB19 GB
Wan 2.2 TI2V 5Bvideo6 GB10 GB
Wan Animate 2 (14B)video12 GB22 GB

VRAM of the graphics card, for the model with the file this site picks for that card. Open a model to see every file and every GPU.

Other versions of this text encoder

FilePrecisionSizeUsed by
umt5_xxl_fp8_e4m3fn_scaled.safetensorsFP86.7 GB12 models
umt5_xxl_fp16.safetensorsFP1611.4 GB12 models

Same encoder, different precision. Any of them works in the same loader node; pick one and select it in the node.

Tested with this file

I use this exact file in 3 of my free workflows, measured on an RTX 5060 Ti 16 GB:

All free workflows →

Not sure your card can run these models? Check your GPU. All shared files: ComfyUI model files.

Size and source read from Hugging Face (2026-10-01).