llamaperf

Qwen3.5 35B (3B active)

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 160,000 ctx

Sep 30, 2026
Throughput
47-51 t/s gen · 1168.0 t/s pp
Quant
UD-IQ3_XXS (GGUF)
KV cache
q4_0
System RAM
128 GB
VRAM reported
16 GB

Summary

User reports Qwen3.5-35B-A3B at 47-51 tok/s generation on an RTX 5060 Ti 16GB over OCuLink in a Proxmox VM. Setup is llama.cpp with UD-IQ3_XXS quant and q4_0 KV cache, fully on the GPU with 348 MiB headroom, at 160K context. Prompt eval of 75K tokens took 64.8 s (1,168 tok/s). The 9B variant reached 40-50 tok/s with UD-Q4_K_XL and about 47-51 tok/s with IQ3_XXS.