llamaperf
Sep 23, 2026
Throughput
53.9 t/s gen · 679.1 t/s pp
Quant
oQ5e
MTP (Multi-Token Prediction)
on
System RAM
128 GB

Use cases

coding

Summary

User benchmarks three Qwen3.8-Flash-Next REAP builds for Apple Silicon on an M4 Max 40-core GPU with 128 GB unified memory at 4K context. Setup is oMLX with mixed-precision packages, 48 transformer layers, top-10 routing, a retained 512-expert native MTP predictor, and 262k configured context. The REAP-384 oQ5e build scored 97/100 HumanEval no-thinking and 61/100 LiveCodeBench no-thinking at 679.1 PP tok/s, 53.9 TG tok/s, and 72.8 GB peak memory. The REAP-384 oQ4e build reached 689.5 PP tok/s and 59.3 TG tok/s at 61.5 GB peak, while REAP-288 oQ4e reached 692.3 PP tok/s and 48.7 TG tok/s at 46.0 GB peak. Accuracy figures are 100-question stratified samples and throughput is workload-specific.