Qwen3.8 125B (6B active) Flash-Next
on M4 Max 128GB · oMLX · 4,096 ctx
Sep 23, 2026
Use cases
coding
Summary
User benchmarks three Qwen3.8-Flash-Next REAP builds for Apple Silicon on an M4 Max 40-core GPU with 128 GB unified memory at 4K context.
Setup is oMLX with mixed-precision packages, 48 transformer layers, top-10 routing, a retained 512-expert native MTP predictor, and 262k configured context. The REAP-384 oQ5e build scored 97/100 HumanEval no-thinking and 61/100 LiveCodeBench no-thinking at 679.1 PP tok/s, 53.9 TG tok/s, and 72.8 GB peak memory.
The REAP-384 oQ4e build reached 689.5 PP tok/s and 59.3 TG tok/s at 61.5 GB peak, while REAP-288 oQ4e reached 692.3 PP tok/s and 48.7 TG tok/s at 46.0 GB peak. Accuracy figures are 100-question stratified samples and throughput is workload-specific.