Qwen3.5 35B (3B active)
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 160,000 ctx
Sep 30, 2026
Summary
User reports Qwen3.5-35B-A3B at 47-51 tok/s generation on an RTX 5060 Ti 16GB over OCuLink in a Proxmox VM.
Setup is llama.cpp with UD-IQ3_XXS quant and q4_0 KV cache, fully on the GPU with 348 MiB headroom, at 160K context.
Prompt eval of 75K tokens took 64.8 s (1,168 tok/s). The 9B variant reached 40-50 tok/s with UD-Q4_K_XL and about 47-51 tok/s with IQ3_XXS.