llamaperf

M6 32GB

APPLE · 32GB unified memory · 2 reports

As of 8 Oct 2026, the models most run on the M6 32GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: LM Studio 1

Run models on your M6 32GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the M6 32GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 32 GB of unified memory.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.
reported speed:
52.2 tokens/s generation
quant:
4-bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma-4-26B-A4B 4-bit at 52.2 tok/s decode on a base Mac mini M6 with 32 GB unified memory. Setup is SwiftLM with MLX 4-bit weights, full GPU offload, peak GPU memory 19.5 GB; the longest prompt that passed was 80.7K tokens at 622 tok/s prefill and 24.3 tok/s decode. Also benchmarks Qwen3.6-35B-A3B 4-bit at 46.7 tok/s (GPU) and 13.2 tok/s (--stream-experts), Qwen3.8-27B 4-bit dense at 9.3 tok/s with 200 tok/s prefill, and Gemma-4-26B-A4B 8-bit at 8.8 tok/s with --stream-experts. An A/B against the prior revision shows 998 tok/s prefill at ~2.3K tokens versus 508 tok/s with swap, and 914 tok/s at ~9.5K tokens where the earlier build aborted on swap.

Oct 5, 2026

Qwen3.6 35B (3B active)

M6 32GB · LM Studio · 32,768 ctx

Tone: positive
reported speed:
63.8 tokens/s generation
quant:
MLX 4-bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

summarizationmultilingualcodingtool-use

User reports Qwen3.6-35B-A3B at 63.8 t/s median decode on a Mac mini M6 with 32 GB unified memory. Setup is LM Studio with the MLX 4-bit build and a 32,768-token context; the MLX build ignored the context setting and loaded 34k to 165k. Reasoning off, max_tokens 600, temperature 0.7, median of 3 runs per prompt. The same model's Splash build (speculative decoding with DFlash2 drafts) reached 51.5-234.1 t/s across the five prompts, slower on German prose and faster on code and JSON. The user recommends Splash where it exists and notes the MLX build was faster than Splash on German prose (63 vs 52-59 t/s).

Oct 4, 2026

Get a weekly email of new M6 32GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Gemma 4 26B (4B active)
M6 32GB
4-bit
SwiftLM
Not reported52.2 tokens/s
Qwen3.6 35B (3B active)
M6 32GB
MLX 4-bit
LM Studio
32,76863.8 tokens/s