llamaperf

Mac M6 LLM performance

Every M6 configuration on llamaperf, M6, with its unified memory, bandwidth and the generation speed people report. Memory decides which models fit; bandwidth decides how fast they run.

3 configurations · 2 reports

As of 8 Oct 2026, the models most run on M6 Macs, each on one configuration, with the median of plain runs (one machine, one request, no speculative decoding):

Thin page (2 of 3 reports needed for indexing). Add yours.

Configurations

ChipMemoryBandwidthMedian t/sReports
M632 GB171 GB/s642
M624 GB171 GB/s0
M616 GB154 GB/s0

Median t/s is plain single-machine runs of Qwen3.6 35B · 3B active at 4-bit, so every row is the same model; a blank means no run of it on that configuration yet. 2 configurations have no report yet.

Models people run on M6 Macs

Latest reports

reported speed:
52.2 tokens/s generation
quant:
4-bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma-4-26B-A4B 4-bit at 52.2 tok/s decode on a base Mac mini M6 with 32 GB unified memory. Setup is SwiftLM with MLX 4-bit weights, full GPU offload, peak GPU memory 19.5 GB; the longest prompt that passed was 80.7K tokens at 622 tok/s prefill and 24.3 tok/s decode. Also benchmarks Qwen3.6-35B-A3B 4-bit at 46.7 tok/s (GPU) and 13.2 tok/s (--stream-experts), Qwen3.8-27B 4-bit dense at 9.3 tok/s with 200 tok/s prefill, and Gemma-4-26B-A4B 8-bit at 8.8 tok/s with --stream-experts. An A/B against the prior revision shows 998 tok/s prefill at ~2.3K tokens versus 508 tok/s with swap, and 914 tok/s at ~9.5K tokens where the earlier build aborted on swap.

Oct 5, 2026

Qwen3.6 35B (3B active)

M6 32GB · LM Studio · 32,768 ctx

Tone: positive
reported speed:
63.8 tokens/s generation
quant:
MLX 4-bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

summarizationmultilingualcodingtool-use

User reports Qwen3.6-35B-A3B at 63.8 t/s median decode on a Mac mini M6 with 32 GB unified memory. Setup is LM Studio with the MLX 4-bit build and a 32,768-token context; the MLX build ignored the context setting and loaded 34k to 165k. Reasoning off, max_tokens 600, temperature 0.7, median of 3 runs per prompt. The same model's Splash build (speculative decoding with DFlash2 drafts) reached 51.5-234.1 t/s across the five prompts, slower on German prose and faster on code and JSON. The user recommends Splash where it exists and notes the MLX build was faster than Splash on German prose (63 vs 52-59 t/s).

Oct 4, 2026

Frequently asked

Which M6 Mac is best for local LLMs?

On Qwen3.6 35B · 3B active at 4-bit, the M6 model the most of these Macs have been measured on, the M6 32GB is the fastest reported, at 63.8 tokens per second (one run). Speed follows the chip tier's memory bandwidth, and which models fit follows the memory size.

How much unified memory do I need on an M6 Mac for local LLMs?

Enough to hold the model's weights at your quant plus its context, with headroom for macOS. Unified memory is shared with the system, so the usable pool is smaller than the number on the box. The VRAM calculator sizes any model against each configuration.

Why is an M6 Mac slower than an NVIDIA card with less memory?

Generation speed is limited by memory bandwidth, not by memory size. M6 configurations on llamaperf range from about 154 to 171 GB/s, while a high-end discrete GPU reads its smaller pool much faster. A Mac wins on which models fit; a discrete card wins on tokens per second for models that fit both.