llamaperf

Qwen3.8 27B Swift-1.5

on M5 Max 64GB · oMLX · 262,000 ctx

Tone: positive
Sep 29, 2026
Throughput
34.6 t/s gen · 400.0 t/s pp
Quant
oQ8e
System RAM
64 GB

Use cases

codingagenticrp

Summary

User reports Swift-1.5-Qwen3.8-27b-oQ8e-mtp at 34.6 tok/s generation and around 400 tok/s prompt processing on an Apple M5 Max 64GB. Setup is oMLX with oQ8e quant, 262k context window, thinking on at xhigh, MTP speculative decoding enabled. User compares against base Qwen3.8 27B (32.8 tok/s, ~400 tok/s prompt processing) and notes Swift 1.5 generates fewer tokens per run (51k vs 77k average) with a similar quality score (85.8 vs 85.0).