| File | llava_llama3_vision.safetensors |
| What it is | LLaVA-Llama3 vision (other) |
| Size | 0.6 GB |
| Folder | ComfyUI/models/clip_vision/ |
| Loader node | Load CLIP Vision (CLIPVisionLoader) |
| Source | Comfy-Org/HunyuanVideo_repackaged on Hugging Face |
Download llava_llama3_vision.safetensors (0.6 GB)
Where to put it
Put the file in ComfyUI/models/clip_vision/ (in the portable version: ComfyUI_windows_portable/ComfyUI/models/clip_vision/). Then press R in ComfyUI, or restart it, so the file shows up in Load CLIP Vision (CLIPVisionLoader).
If a workflow says “Value not in list” for this node, the file name in the workflow is different from yours, or the file is in another folder. Click the node and pick the file from the list.
What it does
A vision encoder reads an input image, for example the start frame in image-to-video. It is only needed in workflows that take an image.
Needed only: I2V only.
Models that use llava_llama3_vision.safetensors
| Model | Type | Fits from | Runs well from |
|---|---|---|---|
| HunyuanVideo (13B, original) | video | 11 GB | 18 GB |
VRAM of the graphics card, for the model with the file this site picks for that card. Open a model to see every file and every GPU.
Not sure your card can run these models? Check your GPU. All shared files: ComfyUI model files.
Size and source read from Hugging Face (2026-10-01).