llamaperf
Oct 3, 2026
Throughput
50.8 t/s gen · 2384.0 t/s pp
Quant
Q8_K_XL (GGUF)
System RAM
128 GB

Summary

User benchmarks Qwen3.6-35B-A3B at 50.8 t/s decode and 2384 t/s prefill on a Strix Halo 128GB using the gufo engine with Q8_K_XL quantization. Setup is gufo with Q8_K_XL GGUF and no KV cache quantization, running on the Ryzen AI Max+ 395 with 128 GB unified memory. The user compares gufo against llama.cpp across multiple quantizations and speculative decoding modes. Gufo achieves higher prefill speeds (up to 2702 t/s on Q6_K_XL) but llama.cpp decodes faster on Q6_K_XL by 3% to 9%. With DFlash2 speculative decoding at 7 draft tokens, gufo reaches 78.6 t/s on Q8_K_XL versus llama.cpp's 51.7 t/s. The user notes that speculative output is not bit-identical to plain greedy output on either engine.