Qwen3.8 27B
M3 Pro 36GB · custom C + Metal runtime
- reported speed:
- 17.7 tokens/s generation
- quant:
- Q4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports a custom C + Metal runtime with speculative decoding at 18 t/s end-to-end on coding tasks and 17.7 t/s for an LRUCache implementation, on 36 GB unified memory. Q4 weights are about 15 GB, mmap'd. Prose runs at roughly 10-11 t/s. TTFT is about 1.4s for a short prompt and about 2.7s for a 128-token prompt. The user compares the custom runtime against llama.cpp and reports it faster.