Qwen3.6 27B
M3 Max 96GB · oMLX
- reported speed:
- 7.4 tokens/s generation · 121.3 tokens/s prompt processing
- quant:
- 8bit (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.6 27B at 7.4 t/s generation and 121.3 t/s prompt processing on an M3 Max 96GB. Setup is oMLX with the MLX 8-bit model at pp1024/tg128, using 28.34 GB peak memory. A pp4096/tg128 run reached 8.8 t/s generation and 133.8 t/s prompt processing. Continuous batching at 4x reached 19.9 t/s aggregate generation. User is new to LLMs and asks whether the slow speed is due to the dense model or a setup problem.