llamaperf

Qwen3.8 27B

on M5 Ultra 256GB · oMLX · 8,192 ctx

Tone: positive
Sep 17, 2026
Throughput
50.0 t/s gen · 1800.0 t/s pp
Quant
q4
MTP (Multi-Token Prediction)
off
System RAM
256 GB

Summary

User reports Qwen3.8 27B at 50 t/s generation and 1800 t/s prompt processing on an M5 Ultra at 8k context. Figures come from the omlx website rather than a local run, with q4 weights and no MTP. The user notes the benchmarks' provenance is unclear but finds them reasonable. User calls the results very promising.