llamaperf
Sep 18, 2026
Throughput
40.0 t/s gen · 500.0 t/s pp
Quant
Q4_K_XL (GGUF)
KV cache
Q8
System RAM
32 GB
VRAM reported
8 GB

Summary

User asks whether speculative decoding (MTP) is still worth enabling for a 35B A3B MoE model offloaded across an 8GB RTX 5060 and 32GB of system RAM. Current llama-server setup uses a Q4_K_XL GGUF with a Q8 KV cache, flash attention on, 4096 batch and ubatch, 16 CPU cores, 40 MoE layers on CPU and 99 GPU layers, yielding about 40 t/s generation and 500 t/s prompt processing without MTP. User recalls earlier reports that MTP hurt prompt processing and wants to know if that is still the case and how others configure llama-server.