llamaperf

Mac M4 LLM performance

Every M4 configuration on llamaperf, M4, M4 Pro and M4 Max, with its unified memory, bandwidth and the generation speed people report. Memory decides which models fit; bandwidth decides how fast they run.

11 configurations · 14 reports

Configurations

ChipMemoryBandwidthMedian t/sFastestReports
M4 Max128 GB546 GB/s25735
M4 Max96 GB546 GB/s--0
M4 Max64 GB546 GB/s24362
M4 Pro64 GB273 GB/s--0
M4 Max48 GB546 GB/s--0
M4 Pro48 GB273 GB/s13204
M4 Max36 GB410 GB/s--0
M432 GB120 GB/s15222
M424 GB120 GB/s--0
M4 Pro24 GB273 GB/s--0
M416 GB120 GB/s20201

Median and fastest are single-machine generation speeds across every model and quant reported, so open a configuration for like-for-like numbers. 6 configurations have no report yet.

Models people run on M4 Macs

Latest reports

Tone: positive
reported speed:
22.0 tokens/s generation
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 8-22 t/s on a 32GB M4 MacBook Air with 21GB of allocations. Setup is a custom inference engine called Cherenkov that combines predictive expert streaming with optional mixed-precision execution, keeping a bounded working set of experts in unified memory rather than loading the entire model. A one-layer lookahead predicts which experts will be needed next and initiates SSD reads.

Sep 11, 2026
reported speed:
1.8 tokens/s generation
quant:
JANG

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks MoE streaming in oMLX on an M4 Pro 48GB across four models. GLM-5.3-Flash-JANG-MTP reaches 1.82 tok/s with 12.36s TTFT, 10.52 GiB after load and 14.68 GiB peak. Qwen3.8-JANG 4S reaches 3.63 tok/s with 8.56s TTFT, 7.04 GiB after load and 11.19 GiB peak. Qwen3.8-JANG 4M reaches 3.19 tok/s with 11.23s TTFT, 7.05 GiB after load and 11.11 GiB peak. DeepSeek-V4-Flash-0731-JANG reaches 2.71 tok/s with 6.52s TTFT, 8.25 GiB after load and 16.87 GiB peak. The primary record uses GLM-5.3-Flash-JANG-MTP. The other models are Qwen3.8 in JANG 4S and 4M quants and DeepSeek V4 Flash in the 0731 variant with the JANG quant. MoE streaming allows larger MoE models to run with a lower memory footprint at the cost of speed.

Sep 10, 2026
reported speed:
25.0 tokens/s generation
quant:
UD-Q2_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running a model in-browser with custom WebGPU kernels, at speed comparable to llama.cpp.

Sep 7, 2026
Tone: positive
reported speed:
18.0 tokens/s generation
quant:
8-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

mathcodingchat

User reports speculative decoding on Apple Silicon, with an 8-bit model going from 8.2 tok/s baseline to 18-26 tok/s. The speedups are 3.27x on math, 2.5x on code and 2.22x on chat, with output byte-identical. The 4-bit model gets about 1.7x at roughly 25 tok/s and needs about 18 GB, while the 8-bit model peaks at about 40 GB and needs a 48 GB Mac. Meta's DFlash numbers on Mac are 1.5x on an M4 Max and 1.8x on an M5 Max on 4-bit.

Sep 7, 2026
Tone: positive
reported speed:
8.2 tokens/s generation
quant:
8-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports mlx-dspark, a speculative decoding project, running an 8-bit model at 8.2 tok/s baseline and 18-26 tok/s with speculative decoding. Speedups are 3.27x on math, 2.5x on code and 2.22x on chat. The 4-bit model reaches about 1.7x at about 25 tok/s and needs about 18 GB. The 8-bit model peaks at about 40 GB and requires a 48 GB Mac. Output is byte-identical to normal decoding.

Sep 7, 2026

Maple Preview 20B (1B active)

M4 16GB · Mference · 131,072 ctx

Tone: positive
reported speed:
20.0 tokens/s generation · 40.0 tokens/s prompt processing
quant:
ternary

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Maple Preview, a 20B A1B MoE model trained from scratch in ternary precision, running on a MacBook Air M4 with 16 GB RAM. Setup is Mference, a fork of turbo-fieldfare, streaming experts from SSD and reducing memory usage to 500-1200 MB. The model has limited world knowledge and multilingual capabilities but can use tools. The user is enthusiastic about the low memory footprint and potential for background agentic workloads.

Sep 7, 2026
Tone: positive
reported speed:
8.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks about an MTP or DFlash head for Qwen 3.8 27B, and mentions 35B-A3B as a daily driver. User reports about 8 tok/s for Qwen 3.6 27B with MTP on 32 GB of unified memory.

Sep 7, 2026
Tone: positive
reported speed:
20.3 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingmath

User reports speculative decoding with mlx-dspark on an 8-bit target at a 2.45x mean speedup, 8.3 to 20.3 tok/s. Setup is mlx-dspark with a 4-bit target at 1.74x and 25.3 tok/s. The 8-bit target with drafter beats plain 4-bit at 14.6 tok/s.

Sep 7, 2026
reported speed:
72.5 tokens/s generation · 1472.4 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6-35B-A3B on a Mac with OMLX, comparing MTP enabled against disabled. MTP shows minimal speedup for the 35B MoE model but roughly 2x for the 27B dense model. Results include pp and tg t/s at various context lengths and batch sizes.

Sep 7, 2026
Tone: positive
reported speed:
11.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash (284B) at about 11 t/s on a 64GB MacBook. The model is 165GB on disk and does not fit in RAM, so it runs via SSD streaming with a 2-bit file. Quality matches official perplexity (6.1250 vs 6.1262). Higher quality settings give 1-2 t/s. An 8GB cache gave 2.04 t/s versus 1.23 t/s with 32GB. Speculative decoding and prefetching did not help.

Sep 7, 2026
Tone: mixed
reported speed:
36.0 tokens/s generation · 517.9 tokens/s prompt processing
quant:
AD-3.84bpw-M64 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a 85 GB GGUF quant of Qwen 3.8 Flash Next running on a 64 GB MacBook with the ngram table offloaded to SSD, at 517.9 t/s prefill and 36 t/s decode. The user notes the quant quality is far from perfect and better versions are planned.

Sep 7, 2026
reported speed:
15.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagentic

User benchmarks Qwen3.8 27B on an M4 Max 128GB across five runtime and quant stacks at contexts from 32K to 256K. The stacks are oMLX AWQ 5-bit, oMLX oQ8e, mlx-dspark, and MTPLX 4-bit and 8-bit. Best decode at 256K is oMLX AWQ 5-bit at 15.2 t/s, while MTPLX collapses at 256K. Prefix caching is crucial for effective speed.

Sep 7, 2026
Tone: positive

User reports GLM 5.3 Flash support in the ds4 branch, running on an M4 Max with 128 GB. No performance numbers are provided.

Aug 28, 2026
Tone: positive
quant:
mixed-4_8bit (mlx)
codingmath

User reports Qwen 3.8 27B as the first model to break 94% on the cupel benchmark. Setup is llama.cpp with the unsloth/Qwen3.8-Flash-Next-GGUF:UD-IQ4_XS quant, which outperformed other 4-bit quants. Qwen 3.8 27B outperformed in coding but lost in general knowledge to Gemma 31B and Qwen 3.6.

Aug 28, 2026

Frequently asked

Which M4 Mac is best for local LLMs?

Among M4 Macs with community reports, the M4 Max 128GB has the highest median generation speed on llamaperf. Speed follows the chip tier's memory bandwidth; which models fit follows the memory size.

How much unified memory do I need on an M4 Mac for local LLMs?

Enough to hold the model's weights at your quant plus its context, with headroom for macOS. Unified memory is shared with the system, so the usable pool is smaller than the number on the box. The VRAM calculator sizes any model against each configuration.

Why is an M4 Mac slower than an NVIDIA card with less memory?

Generation speed is limited by memory bandwidth, not by memory size. M4 configurations on llamaperf range from about 120 to 546 GB/s, while a high-end discrete GPU reads its smaller pool much faster. A Mac wins on which models fit; a discrete card wins on tokens per second for models that fit both.