- reported speed:
- 51.9 tokens/s generation · 322.0 tokens/s prompt processing
- quant:
- Q4_K_M (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Dual GPU setup with RX 7900 GRE 16GB and RX 480 8GB. Benchmarks show 36% improvement in generation speed with dual GPU. Also tested medgemma-27b-it-UD-Q6_K_XL, Qwen3.8-27B-Q6_K, Qwen3.8-27B-OBLITERATED-Q5_K_M, and Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M with and without Flash Attention.