Datasheet for local AI50 models99 GPUsData read 2026-10-01
No adsNo tracking
Check my GPU →

RTX 5060 Ti 16 GB for local AI

Which image and video models run on the RTX 5060 Ti 16 GB, which file to download for each, and how much VRAM they need.

NVIDIA16 GB GDDR7 · 128-bit · 448 GB/s · BlackwellData 2026-10-01
VRAM16 GB
MemoryGDDR7
Bus128-bit
Bandwidth448 GB/s
ArchitectureBlackwell
Launched2025-04
Launch price$429
Runs well26 of 50

Specs: www.nvidia.com · launch: www.nvidia.com

01What runs on it

ModelVerdictBest fileSizeNeeded
image models
FLUX.1 [dev]Runs wellFP811.9 GB14.2 GB
FLUX.1 [schnell]Runs wellFP811.9 GB14.2 GB
FLUX.1 Kontext [dev]Runs wellFP811.9 GB14.5 GB
FLUX.1 Krea [dev]Runs wellFP811.9 GB14.2 GB
FLUX.1 Fill [dev]Runs wellQ8_012.7 GB15.3 GB
FLUX.2 [dev]Offload onlyQ2_K12.9 GB16.2 GB
FLUX.2 [klein] 9BRuns wellFP89.4 GB11.7 GB
FLUX.2 [klein] 4BRuns well16-bit7.8 GB9.6 GB
Krea 2 (Turbo)Runs wellFP813.1 GB15.7 GB
Qwen-ImageRunsQ4_K_M13.1 GB15.9 GB
Qwen-Image-Edit (2511)TightQ3_K_M9.9 GB12.9 GB
Qwen-Image 2.1Runs wellINT87.3 GB9.6 GB
Z-Image TurboRuns well16-bit12.3 GB14.3 GB
Z-Image (base)Runs well16-bit12.3 GB14.3 GB
Ideogram 4RunsQ4_16.2 GB14.7 GB
Boogu-Image (Turbo)Runs wellFP810.3 GB12.9 GB
ERNIE-Image (Turbo)Runs wellQ8_08.7 GB11.0 GB
HiDream-O1-ImageRuns wellFP88.1 GB10.9 GB
Mage-Flow (Microsoft)Runs well16-bit8.2 GB10.0 GB
Ming-Image 0.1 DesignRuns well16-bit12.3 GB14.6 GB
Lumina Image 2.0Runs well16-bit5.2 GB7.0 GB
HiDream-I1 (Full)RunsQ5_K_M13.0 GB15.8 GB
HiDream-I1 (Dev)RunsQ5_K_M13.0 GB15.8 GB
Stable Diffusion 3.5 LargeRuns wellQ8_08.8 GB11.1 GB
Stable Diffusion 3.5 MediumRuns well16-bit5.1 GB6.9 GB
Chroma1-HDRuns wellFP89.2 GB11.5 GB
SDXL 1.0Runs well16-bit6.9 GB7.1 GB
Illustrious XL / Pony (SDXL anime)Runs well16-bit6.9 GB7.1 GB
Stable Diffusion 1.5Runs well16-bit2.1 GB3.3 GB
HunyuanImage 2.1RunsQ4_K_M11.3 GB14.6 GB
video models
Wan 2.1 T2V 14BRunsQ5_K_M11.3 GB15.6 GB
Wan 2.1 T2V 1.3BRuns well16-bit2.8 GB5.6 GB
Wan 2.1 I2V 14B 480PRunsQ4_K_M11.3 GB15.6 GB
Wan 2.1 I2V 14B 720PTightQ3_K_M8.6 GB15.4 GB
Wan 2.1 VACE 14BTightQ3_K_S7.8 GB12.6 GB
Wan 2.2 T2V A14BRunsQ5_K_M10.8 GB15.1 GB
Wan 2.2 I2V A14BRunsQ5_K_M10.8 GB15.1 GB
Wan 2.2 TI2V 5BRuns well16-bit10.0 GB13.8 GB
Wan 2.2 Animate 14BTightQ3_K_M8.6 GB13.9 GB
Wan Animate 2 (14B)TightQ3_K_M8.6 GB13.9 GB
Wan 2.2 S2V 14BTightQ2_K9.5 GB14.3 GB
SCAIL-2 (character animation)TightQ3_K_M9.1 GB14.4 GB
HunyuanVideo (13B, original)RunsQ6_K11.0 GB15.3 GB
LTX-Video 13B (0.9.8)RunsQ6_K10.9 GB15.2 GB
HunyuanVideo 1.5Runs wellFP88.3 GB12.6 GB
LTX-2 (19B)TightQ3_K_M10.1 GB14.9 GB
LTX-2.3 (22B)TightQ3_K_M10.8 GB15.6 GB
LTX-2.5 (22B)TightQ2_K8.8 GB13.6 GB
MiniMax H3 (33B)Offload onlyQ4_K_M19.9 GB25.7 GB
MiniMax H3 PrunedTightQ3_K_M8.9 GB14.7 GB

Calculated from real file sizes plus working memory. How the numbers work.

02Good to know

The RTX 5060 Ti 16 GB is a Blackwell GPU with FP8 and FP4 hardware: ComfyUI computes Comfy-Org's FP8 files natively here, and NVFP4 files (where a model offers them) are faster still.

03Measured and reported results

LabelModelSetupResultPeak VRAMDateSource
measuredFLUX.1 devflux1-dev-Q8_0.gguf · 1024x1024 · 20 steps
3 timed runs after 1 warm-up (50.7 s incl. loading). T5-XXL FP8 + CLIP-L, euler/simple, guidance 3.5, ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Time = sampling + VAE
40.62 s / image15.1 GB2026-09-27
measuredFLUX.1 devflux1-dev-Q4_K_S.gguf · 1024x1024 · 20 steps
3 timed runs after 1 warm-up (52.1 s incl. loading). T5-XXL FP8 + CLIP-L, euler/simple, guidance 3.5, ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Time = sampling + VAE
47.28 s / image8.8 GB2026-09-27
measuredFLUX.1 devflux1-dev-fp8.safetensors · 1024x1024 · 20 steps
3 timed runs after 1 warm-up (39.1 s incl. loading). T5-XXL FP8 + CLIP-L, euler/simple, guidance 3.5, ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Time = sampling + VAE
36.36 s / image14.8 GB2026-09-27
measuredFLUX.1 Kontextflux1-dev-kontext_fp8_scaled.safetensors · 1024x1024 · 20 steps
Image editing: one 1024×1024 input image, 20 steps, euler/simple, guidance 2.5, ReferenceLatent + FluxKontextImageScale. 2 timed runs after 1 warm-up (60.0 s incl. loading). Editin
53.42 s / image15.0 GB2026-10-11
measuredFLUX.1 Kontextflux1-kontext-dev-Q8_0.gguf · 1024x1024 · 20 steps
Image editing, same settings as the FP8 run: one 1024×1024 input image, 20 steps, euler/simple, guidance 2.5. 2 timed runs after 1 warm-up (94.1 s incl. loading). The GGUF file (Co
82.06 s / image15.5 GB2026-10-11
measuredFLUX.1 schnellflux1-schnell-Q8_0.gguf · 1024x1024 · 4 steps
3 timed runs after 1 warm-up (22.8 s incl. loading). T5-XXL FP8 + CLIP-L, euler/simple, ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Time = sampling + VAE decode. Peak V
9.42 s / image15.1 GB2026-09-27
measuredFLUX.2 klein 4Bflux-2-klein-4b-fp8.safetensors · 1024x1024 · 4 steps
3 timed runs after 1 warm-up (6.66 s incl. loading). Text encoder Qwen3 4B FP8. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. 1024x1024, 4 steps, euler, Flux2Scheduler, C
2.85 s / image12.2 GB2026-10-01
measuredFLUX.2 klein 4Bflux-2-klein-4b.safetensors · 1024x1024 · 4 steps
3 timed runs after 1 warm-up (6.94 s incl. loading). Text encoder Qwen3 4B FP8. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. 1024x1024, 4 steps, euler, Flux2Scheduler, C
3.93 s / image14.1 GB2026-10-01
measuredFLUX.2 klein 4Bflux-2-klein-4b-Q8_0.gguf · 1024x1024 · 4 steps
3 timed runs after 1 warm-up (8.95 s incl. loading). Text encoder Qwen3 4B FP8. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. 1024x1024, 4 steps, euler, Flux2Scheduler, C
4.34 s / image12.4 GB2026-10-01
measuredFLUX.2 klein 9Bflux-2-klein-9b-Q8_0.gguf · 1024x1024 · 4 steps
3 timed runs after 1 warm-up (19.59 s incl. loading). Text encoder Qwen3 8B FP8. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. 1024x1024, 4 steps, euler, Flux2Scheduler,
8.52 s / image13.5 GB2026-10-01
measuredFLUX.2 klein 9Bflux-2-klein-9b-Q4_K_M.gguf · 1024x1024 · 4 steps
3 timed runs after 1 warm-up (17.11 s incl. loading). Text encoder Qwen3 8B FP8. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. 1024x1024, 4 steps, euler, Flux2Scheduler,
9.85 s / image13.5 GB2026-10-01
measuredKrea 2krea2_turbo_int8_convrot.safetensors · 1024x1024 · 8 steps
3 timed runs after 1 warm-up (15.56 s incl. loading). Text encoder Qwen3-VL 4B FP8 (qwen3vl_4b_fp8_scaled), Qwen-Image VAE. 8 steps, euler/simple, CFG 1. ComfyUI 0.38.1 portable, P
10.95 s / image15.3 GB2026-10-04
measuredKrea 2krea2_turbo_fp8_scaled.safetensors · 1024x1024 · 8 steps
3 timed runs after 1 warm-up (22.59 s incl. loading). Text encoder Qwen3-VL 4B FP8 (qwen3vl_4b_fp8_scaled), Qwen-Image VAE. 8 steps, euler/simple, CFG 1. ComfyUI 0.38.1 portable, P
17.74 s / image14.9 GB2026-10-04
measuredMing-Imageming_image_0.1_design_int8_convrot.safetensors · 1024x1024 · 12 steps
2 timed runs after 1 warm-up (17.57 s incl. loading). Design variant. Text encoder Ling-Mini-2.0 INT8 (ming_image_0.1_ling_mini_2.0_int8_convrot), Ming-Image VAE. 12 steps, euler/s
7.42 s / image13.6 GB2026-10-04
measuredMing-Imageming_image_0.1_design_int8_convrot.safetensors · 2048x2048 · 12 steps
2 timed runs after 1 warm-up (52.11 s incl. loading). Design variant. Text encoder Ling-Mini-2.0 INT8 (ming_image_0.1_ling_mini_2.0_int8_convrot), Ming-Image VAE. 12 steps, euler/s
44.85 s / image14.4 GB2026-10-04
measuredMing-Imageming_image_0.1_design_bf16.safetensors · 1024x1024 · 12 steps
2 timed runs after 1 warm-up (28.3 s incl. loading). Design variant. Text encoder Ling-Mini-2.0 INT8 (ming_image_0.1_ling_mini_2.0_int8_convrot), Ming-Image VAE. 12 steps, euler/si
17.39 s / image15.3 GB2026-10-04
measuredQwen-Image 2.1qwen_image_2.1_int8_convrot.safetensors · 1024x1024 · 25 steps
2 timed runs after 1 warm-up (24.6 s incl. loading). Text encoder Qwen3-VL 8B INT8 (qwen3vl_8b_int8_convrot), Qwen-Image 2.1 VAE. 25 steps, euler/simple, CFG 1, TextEncodeQwenImage
17.16 s / image14.3 GB2026-10-04
measuredQwen-Image 2.1qwen_image_2.1_Q8_0.gguf · 1024x1024 · 25 steps
2 timed runs after 1 warm-up (53.91 s incl. loading). Text encoder Qwen3-VL 8B INT8 (qwen3vl_8b_int8_convrot), Qwen-Image 2.1 VAE. 25 steps, euler/simple, CFG 1, TextEncodeQwenImag
43.57 s / image14.8 GB2026-10-04
measuredQwen-Image 2.1qwen_image_2.1_bf16.safetensors · 1024x1024 · 25 steps
2 timed runs after 1 warm-up (46.2 s incl. loading). Text encoder Qwen3-VL 8B INT8 (qwen3vl_8b_int8_convrot), Qwen-Image 2.1 VAE. 25 steps, euler/simple, CFG 1, TextEncodeQwenImage
39.67 s / image15.5 GB2026-10-04
measuredQwen-Image-Editqwen-image-edit-2511-Q4_K_M.gguf · 1024x1024 · 4 steps
Image editing with the Lightning 4-step LoRA (Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16), cfg 1, euler/simple, ModelSamplingAuraFlow 3.1, CFGNorm, index_timestep_zero referen
27.19 s / image15.9 GB2026-10-11
measuredWan 2.2 5Bwan2.2_ti2v_5B_fp16.safetensors · 1280x704 · 20 steps
Text to video, 121 frames (about 5 s at 24 fps). 2 timed runs (516.2 and 514.0 s) after 1 warm-up (528.9 s incl. loading). UMT5-XXL FP8 + Wan 2.2 VAE, uni_pc/simple, CFG 5, shift 8
515.1 s / clip (121 frames)11.0 GB2026-09-27
measuredWan 2.2 T2Vwan2.2_t2v_high/low_noise_14B_fp8_scaled.safetensors · 832x480 · 20 steps
1 timed run after 1 warm-up (1149.4 s). Peak RAM 28.6 GB. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Text to video, 81 frames (about 5 s), 832x480, 20 steps (10 high-n
1146.1 s / clip (81 frames)13.3 GB2026-09-27
measuredWan 2.2 T2VWan2.2-T2V-A14B-HighNoise/LowNoise-Q5_K_M.gguf · 832x480 · 20 steps
2 timed runs (1535.0, 1544.5 s) after 1 warm-up (1538.7 s). Peak RAM 33.1 GB. ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Text to video, 81 frames (about 5 s), 832x480,
1539.8 s / clip (81 frames)14.0 GB2026-09-27
measuredWan 2.2 T2VWan2.2-T2V-A14B-HighNoise/LowNoise-Q8_0.gguf · 832x480 · 20 steps
1 timed run after 1 warm-up (1516.6 s). Peak RAM 32.9 GB (warm-up), 31.9 GB (timed run). ComfyUI 0.25.0 portable, PyTorch 2.12 cu130, 32 GB RAM. Text to video, 81 frames (about 5 s
1515.9 s / clip (81 frames)13.1 GB2026-09-27
measuredZ-Image Turboz_image_turbo_bf16.safetensors · 1024x1024 · 8 steps
3 timed runs after 1 warm-up (15.4 s incl. loading). Qwen3 4B FP8 text encoder + ae VAE, res_multistep/simple, CFG 1, shift 3. Time = sampling + VAE decode. ComfyUI 0.25.0 portable
11 s / image14.6 GB2026-09-27
reportedFLUX.1 devfp8 (ComfyUI template)
“Prompt executed in 25.71 seconds”
ComfyUI 'GPU Benchmark Flux DEV fp8' thread: stock Flux dev fp8 workflow template, time of 2nd/3rd run (no loading); resolution/steps not stated in thread; poster's log shows 16311
25.71 s / image—2025-08-04github.com →
reportedFLUX.2 devflux2_dev_fp8mixed + mistral_3_small_flux2_fp8 · 1024x1024 · 20 steps
“FP8 t2i 142秒程度(20Steps)”
ComfyUI t2i at 1MP; approximate (程度)
142 s / image—2025-11-28note.com →
reportedFLUX.2 devGGUF Q4_0 (flux2-dev-q4_0.gguf) · 1024x1024 · 20 steps
“GGUF Q4_0 t2i 180秒程度(20Steps)”
ComfyUI t2i at 1MP; approximate (程度); author notes GGUF slower than fp8 when spilling VRAM
180 s / image—2025-11-28note.com →
reportedIllustrious / PonyIllustrious-XL-v2.0 · 1024x1024
“NVIDIA RTX 5060 Ti | 2.60it/s | 0.39s/it | CUDA 12.9 | ComfyUI (Unknown) | Arch Linux”
Community speed list; ComfyUI, Euler/Normal, CFG 8, square 1:1, batch speed only (no per-image time); date = list last-updated date; SDXL 1024px section; row says 'RTX 5060 Ti', co
0.39 s/it—2025-11-29huggingface.co →
reportedWan 2.2 I2Vvideo_wan2_2_14B_i2v template (quant not stated) · 720p
“RTX5060ti 16GBだと165.9秒なので、2〜3倍時間が必要ですが、落ちずに720p動画が生成できる事はたいしたものです。”
Comparison figure given in the RTX 3060 post, presumably same 720p/53-frame test (not explicitly restated); steps not stated
165.9 s / clip (53 frames)—2025-11-26note.com →
reportedZ-Image Turboz_image_turbo_bf16.safetensors · 1328x1328 · 8 steps
“約 35秒かかりました。一度モデルを VRAM にロードした後、プロンプトを変えての再実行だと約 22秒で生成できます。”
ComfyUI; ~35 s first run incl. load, ~22 s warm; approximate ('約')
22 s / image—2025-11-30iwannacreateapps.com →

Reported results are other people's numbers, copied as published, with a link. Settings, drivers and ComfyUI versions differ, so compare them with care. Send yours.