- reported speed:
- 20.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at around 20 t/s sustained on an M5 Pro 64GB, with about 30 t/s in the first 2-4k tokens. User finds it unbearably slow, possibly because the model thinks so much. User also tried antirez's DS4 with DeepSeek V4 Flash and gets around 8 t/s sustained, which they call slow and basically unusable.