Qwen3.6 35B (3B active)
on NVIDIA RTX 5060 8GB · llama.cpp
Sep 18, 2026
Summary
User asks whether speculative decoding (MTP) is still worth enabling for a 35B A3B MoE model offloaded across an 8GB RTX 5060 and 32GB of system RAM.
Current llama-server setup uses a Q4_K_XL GGUF with a Q8 KV cache, flash attention on, 4096 batch and ubatch, 16 CPU cores, 40 MoE layers on CPU and 99 GPU layers, yielding about 40 t/s generation and 500 t/s prompt processing without MTP.
User recalls earlier reports that MTP hurt prompt processing and wants to know if that is still the case and how others configure llama-server.