Qwen3.6 35B (3B active)
on M6 32GB · LM Studio · 32,768 ctx
Oct 4, 2026
Use cases
summarizationmultilingualcodingtool-use
Summary
User reports Qwen3.6-35B-A3B at 63.8 t/s median decode on a Mac mini M6 with 32 GB unified memory.
Setup is LM Studio with the MLX 4-bit build and a 32,768-token context; the MLX build ignored the context setting and loaded 34k to 165k. Reasoning off, max_tokens 600, temperature 0.7, median of 3 runs per prompt.
The same model's Splash build (speculative decoding with DFlash2 drafts) reached 51.5-234.1 t/s across the five prompts, slower on German prose and faster on code and JSON. The user recommends Splash where it exists and notes the MLX build was faster than Splash on German prose (63 vs 52-59 t/s).