llamaperf

Qwen3.5 35B (3B active)

on M3 Max 128GB · oMLX

Tone: positive
Oct 7, 2026
Throughput
90.8 t/s gen
Quant
4-bit (MLX)
System RAM
128 GB

Use cases

agenticsummarizationcreative-writinglong-context

Summary

User benchmarks Qwen3.5-35B-A3B (thinking disabled) on Apple Silicon, reporting effective throughput of 71.3 tok/s on the ops-agent scenario with an M3 Max 128GB (40 GPU) running oMLX at MLX 4-bit; generation speed was 90.8 tok/s. Setup is oMLX with a tiered KV cache; the same M3 Max run gives 61.4 effective tok/s on doc-summary, 22.6 on prefill-test and 90.1 on creative-writing. Generation speed is similar across MLX engines (~55-93 tok/s depending on hardware) but prefill varies widely: at 8K context LM Studio MLX takes 49s to prefill while oMLX takes 1.7s with its persistent SSD cache from a prior run. Effective tok/s is output tokens divided by total wall-clock time including prefill. On the same M3 Max, LM Studio MLX reaches 37.1 effective tok/s; on M1 Max 64GB (24 GPU) oMLX 4-bit reaches 37.5 and oMLX fp16 47.3.