llamaperf

M3 Max 48GB

APPLE · 48GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 48 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M3 Macs compared →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 Flash-Next

M3 Max 48GB · oMLX · 65,536 ctx

Tone: mixed
reported speed:
38.0 tokens/s generation
quant:
T5 (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingmath

User reports a custom oMLX fork with a Metal kernel for ternary experts running at 30.7 t/s on a 65,536-token prompt plus 256 output tokens, with prefill at 210-380 tok/s. Setup is a ternary routed expert gate/up (Bonsai T5 packing, ~1.875 bpw), Q3 expert down projections, and a 53 GB n-gram table left on SSD as Q8 and mmapped per token, all low-bit tensors fitted with Unsloth's imatrix. Stock oMLX/mlx-lm will not load it. 64K context is confirmed; 96K trips the prefill guard. Physical peak is 42.3 GiB at 64K and ~41.5 GiB at 8K, with ~35.6 GiB resident and ~2 GiB swapped once at load. The oMLX memory guard is on the 'safe' profile with a 48 GB limit, one model and one request at a time. Against Unsloth UD-Q4_K_XL, which does not fit in 48 GB, the user reports KLD vs Q8_0 of 0.49 vs 0.036, MMLU 83.0% vs 89.7%, GSM8K 90.0% vs 92.0%, and HumanEval 92.7% vs 95.7%. The model sometimes ignores 'answer with just the letter' in Chinese.

Sep 10, 2026