llamaperf

Qwen3.6 35B (3B active)

on NVIDIA RTX 4060 · llama.cpp

Tone: mixedNVIDIA hardware
Oct 8, 2026
Throughput
24.2 t/s gen · 65.6 tokens/s prompt processing (prefill)Prompt processing measures input prefill speed; it does not tell you the time to first token.
Quant
IQ4_XS (GGUF)
System RAM
32 GB
VRAM reported
8 GB

Summary

User reports Qwen3.6-35B-A3B at 24.2 t/s generation and 65.62 t/s prompt eval on an RTX 4060 8GB with 32GB DDR5 RAM. Setup is llama.cpp with IQ4_XS quant and --moe-cache-mib 2048, using the PR#29887 MoE expert cache in host memory. User is not getting expected t/s and asks for help; also tried Qwen3.8-Flash-Next Q2 at 8.02 t/s generation and 6.96 t/s prompt eval.