Oct 3, 2026
Summary
User benchmarks Qwen3.6-35B-A3B at 50.8 t/s decode and 2384 t/s prefill on a Strix Halo 128GB using the gufo engine with Q8_K_XL quantization.
Setup is gufo with Q8_K_XL GGUF and no KV cache quantization, running on the Ryzen AI Max+ 395 with 128 GB unified memory. The user compares gufo against llama.cpp across multiple quantizations and speculative decoding modes.
Gufo achieves higher prefill speeds (up to 2702 t/s on Q6_K_XL) but llama.cpp decodes faster on Q6_K_XL by 3% to 9%. With DFlash2 speculative decoding at 7 draft tokens, gufo reaches 78.6 t/s on Q8_K_XL versus llama.cpp's 51.7 t/s. The user notes that speculative output is not bit-identical to plain greedy output on either engine.