llamaperf
Sep 22, 2026
Throughput
70.0 t/s gen
Quant
Q4_K_XL (GGUF)
System RAM
16 GB
VRAM reported
16 GB

Summary

User reports almost 70 t/s with Qwen3.6-35B-A3B Q4_K_XL on a budget build using three Nvidia Tesla P100 16GB GPUs. The GPUs cost $80 each and are split across three nodes to manage thermals; the motherboard required a patched BIOS to enable Above 4G Decoding. User is still testing and optimizing, and has published a GitHub repository for the build.