Qwen3.8 125B (6B active) Flash-Next
M4 32GB · Cherenkov
- reported speed:
- 22.0 tokens/s generation
- quant:
- Q4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 8-22 t/s on a 32GB M4 MacBook Air with 21GB of allocations. Setup is a custom inference engine called Cherenkov that combines predictive expert streaming with optional mixed-precision execution, keeping a bounded working set of experts in unified memory rather than loading the entire model. A one-layer lookahead predicts which experts will be needed next and initiates SSD reads.