- reported speed:
- 2.7 tokens/s generation
- quant:
- Q4_K_M (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
User benchmarked Qwen3.6-27B Q4_K_M locally on RX 5700 XT 8GB, achieving 2.70 tok/s. Also tested other models including Qwen3.5, Qwen3.6-35B-A3B, and Gemma-4-31b-it. Subjective ranking placed Qwen3.6-27B Q4_K_M second overall for the coding task.