Qwen3.8 27B Swift-1.5
M5 Ultra 96GB · mlx-serve · 25,000 ctx
- reported speed:
- 113.7 tokens/s generation · 3191.0 tokens/s prompt processing
- quant:
- 4.7bpw (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 Swift-1.5 at 113.7 tok/s decode and 3191 tok/s prefill on an M5 Ultra 96GB Mac Studio. Setup is mlx-serve 26.10.1 with a 4.7bpw MLX quantization, 107GB download, text-only, tested to 179,200 tokens of context. Decode falls to 81.1 tok/s after a 95k prompt, where prefill is 2928 tok/s. User compares against a llama.cpp IQ3_XXS build of the same model, which reaches 62.7 tok/s decode after 4k and 1427 tok/s prefill at 25k. Top-1 agreement with Swift BF16 is 91.0% across 680 held-out positions, against 84.1% for the llama.cpp build.