llamaperf
Sep 18, 2026
Throughput
19.7 t/s gen · 332.6 t/s pp
Quant
MXFP4 (GGUF)
System RAM
60 GB

Summary

User benchmarks Qwen3.6 35B-A3B at 19.7 t/s generation and 332.57 t/s prompt processing on a mixed GTX 1080 Ti and Radeon VII Vulkan setup. Setup is llama.cpp Vulkan build b29c606e2 with the MXFP4 MoE GGUF and Flash Attention enabled, across two GPUs. User also reports results for twelve other models including llama 7B, bailingmoe2 16B-A1B, gpt-oss 20B, Gemma 4 26B-A4B, Qwen3.5 27B, Qwen3-Coder 30B-A3B, granite 4.0, and Phi-3.5-MoE.