Qwen3.8 27B
M1 Max 32GB · MLX
- reported speed:
- 15.8 tokens/s generation · 81.8 tokens/s prompt processing
- quant:
- 4bit (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User compares MLX and llama.cpp on Qwen3.8-27B. MLX uses mlx-community/Qwen3.8-27B-4bit at ~16.1GB, with prompt processing at 81.76 tok/s, generation at 15.81 tok/s, and peak memory of 16.39GB. llama.cpp uses unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M at 15.32 GiB and 27.32B params, with full Metal offload and Flash Attention enabled, giving prompt processing at 99.61 ± 0.44 tok/s and generation at 9.69 ± 0.34 tok/s. llama.cpp is ~22% faster at prompt processing, while MLX is ~63% faster at generation.