Qwen3.8 27B
M2 Pro 32GB · llama.cpp · 131,072 ctx
- reported speed:
- 8.6 tokens/s generation · 21.9 tokens/s prompt processing
- quant:
- IQ4_XS (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingvision
User reports Qwen3.8 27B IQ4_XS with vision projector on an M2 MacBook Pro 32GB, at 21.9 t/s prompt and 8.6 t/s generation. Setup is llama.cpp built from source, with context set to 128K due to memory constraints. User notes the model is slow on Mac but that thinking quality is good, and mentions previous experience with Qwen3.5 35B-A3B.