Datasheet for local AI50 models99 GPUsData read 2026-10-01
No adsNo tracking
Check my GPU →

llava_llama3_vision.safetensors

LLaVA-Llama3 vision — the vision encoder for HunyuanVideo 13B. 0.6 GB. In ComfyUI it goes in models/clip_vision.

0.6 GBmodels/clip_vision1 model
Filellava_llama3_vision.safetensors
What it isLLaVA-Llama3 vision (other)
Size0.6 GB
FolderComfyUI/models/clip_vision/
Loader nodeLoad CLIP Vision (CLIPVisionLoader)
SourceComfy-Org/HunyuanVideo_repackaged on Hugging Face

Download llava_llama3_vision.safetensors (0.6 GB)

Where to put it

Put the file in ComfyUI/models/clip_vision/ (in the portable version: ComfyUI_windows_portable/ComfyUI/models/clip_vision/). Then press R in ComfyUI, or restart it, so the file shows up in Load CLIP Vision (CLIPVisionLoader).

If a workflow says “Value not in list” for this node, the file name in the workflow is different from yours, or the file is in another folder. Click the node and pick the file from the list.

What it does

A vision encoder reads an input image, for example the start frame in image-to-video. It is only needed in workflows that take an image.

Needed only: I2V only.

Models that use llava_llama3_vision.safetensors

ModelTypeFits fromRuns well from
HunyuanVideo (13B, original)video11 GB18 GB

VRAM of the graphics card, for the model with the file this site picks for that card. Open a model to see every file and every GPU.

Not sure your card can run these models? Check your GPU. All shared files: ComfyUI model files.

Size and source read from Hugging Face (2026-10-01).