llamaperf
Sep 22, 2026
Throughput
49.1 t/s gen · 1238.4 t/s pp
MTP (Multi-Token Prediction)
on
System RAM
256 GB

Summary

User reports Mimo 2.6 Flash at 49.1 t/s generation and 1238.4 t/s prompt processing on an M5 Ultra 256GB, averaged over five trials at 32,000 prompt tokens and 64 generation tokens. Setup is a custom mlx-vlm patch loading the original weights, with MTP enabled but apparently not working. A second run at 64,000 prompt tokens and 1024 generation tokens averaged 38.1 t/s generation and 981.9 t/s prompt processing. User also compares Qwen3.8 Flash-Next FP8 on oMLX, reaching 39.7 to 43.9 t/s generation across 32,768 to 200,000 token prompts.