llamaperf

M2 Ultra 192GB

APPLE · 192GB unified memory · 3 reports

See what fits on this GPU →

Use the calculator to check which models fit in 192 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M2 Macs compared →

Qwen3.8 27B

M2 Ultra 192GB · llama.cpp · 131,072 ctx

reported speed:
22.4 tokens/s generation · 360.2 tokens/s prompt processing
quant:
Q6_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports llama-bench results and a serving setup, with about 16 t/s on WebUI with a specific prompt.

Sep 7, 2026
Tone: positive
reported speed:
25.8 tokens/s generation · 250.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a custom llama.cpp fork on an M2 Ultra with a repacked model at 141 GiB, smaller than Q4 GGUF, peaking at 42 t/s generation. Setup uses an SSD KV cache and dynamic lanes, 8 lanes. Prompt processing runs about 250 t/s at 8k-32k context.

Sep 7, 2026
reported speed:
28.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports decode speeds at various context depths: 28 t/s at the start, 23.5 t/s at 45k, and 18 t/s at 192k. The run was maintained with 8k token output. Prefill performance is mentioned but no numbers are given.

Aug 3, 2026