llamaperf
← All articles

Mac setups

M5 Ultra Mac Studio for local LLMs: early measurements

By llamaperf · · 7 min read

Quick answer

The M5 Ultra Mac Studio writes text about 1.6 times as fast as an M5 Max in llama.cpp's standard test, and reads prompts far faster than the M3 Ultra it replaces. An RTX 5090 still writes faster when the model fits in its 32 GB. Buy the Ultra for memory and long prompts, and size it against the share macOS lets the GPU use.

What Apple shipped

Four chips, and the number that matters

The new Mac Studio went on sale on 22 September with two chips in four configurations. For a local model the figure to read first is memory bandwidth, because writing each token means reading the model's weights out of memory again. Here's what Apple's tech specs list.

ChipGPU coresMemoryBandwidth
M5 Max3236 GB460 GB/s
M5 Max4048, 64 or 128 GB614 GB/s
M5 Ultra6496 GB1.2 TB/s
M5 Ultra80256 or 512 GB1.2 TB/s

The 96 GB Ultra is the cut-down 64-core part. People in threads sometimes assume it's the full chip, and it isn't. The 256 GB and 512 GB models come only with the 80-core GPU, and Apple says the 512 GB version arrives later than the rest.

What Apple's own claim covers

Apple's launch release makes one LLM claim:

"Up to 9.8x faster LLM prompt processing in LM Studio when compared to Mac Studio with M1 Ultra, and up to 4x faster than M3 Ultra." (Apple Newsroom, 25 August 2026)

That's prompt processing, the part where the model reads what you sent. It says nothing about how fast text comes back, and the release names no model or quantization. So treat it as a ceiling for long prompts and look at owners' numbers for the rest.

What owners measured

The same test on three Macs

The most useful comparison is the long-running Apple Silicon thread in llama.cpp, where everyone runs the same small model through llama-bench with speculation off and posts the result (llama.cpp discussion 4167). These are the rows for its 8-bit version.

ChipReads the prompt (t/s)Writes text (t/s)
M5 Ultra, 80 cores4,910.55119.01
M5 Max, 40 cores3,143.8172.42
M3 Ultra, 80 cores1,487.5163.93

The M5 Ultra row was posted on 30 September by an owner of a 256 GB Mac Studio. As of 1 October nobody had posted the 64-core Ultra or the 32-core Max, so the 96 GB and 36 GB models were still unmeasured.

A real model at a real prompt length

Small test models flatter every chip. MacStories ran Qwen3.8 27B on a 6,091-token prompt with speculation off, and the M5 Ultra read it at 1,701 tokens a second and wrote at 48. The M3 Ultra managed 414 and 31. An RTX 5090 in the same test read at 3,031 and wrote at 59 (MacStories review). Federico Viticci explains the gap in plain terms:

"Token generation, on the other hand, is bandwidth: the model writes one token at a time and pulls the entire model back out of memory for each one" (MacStories)

The 5090 has 1.79 TB/s against the Ultra's 1.2, so it stays ahead on writing as long as the model fits in its 32 GB. Past that, the Mac wins by default, because the 5090 has to borrow system memory. Reports from people running these machines, with each run's model, quant and engine, are on the Mac page.

Ultra or Max

Twice the bandwidth isn't twice the speed

On paper the Ultra has double the Max's bandwidth. In the llama.cpp test it writes 1.64 times as fast (119.01 against 72.42). Every Ultra so far has lost some speed to the link between its two halves, and the M5 Ultra does too. If your models fit in 128 GB and you mostly chat, an M5 Max gets you most of the way for less than half the starting price ($2,499 against $5,499 in the US, per Apple's release).

The Ultra earns its price in two places. Long prompts read much faster, which matters if you paste in whole codebases or long documents. And it's the only way to get more than 128 GB on a Mac. If you're unsure which camp you're in, you're probably a Max buyer: the people who need an Ultra tend to know exactly which model they want to run and how much memory it takes.

The memory a model can use

macOS won't give the GPU all of your memory. By default it allows about three quarters on Macs with 36 GB or more, so a 96 GB Ultra gives a model about 72 GB and a 256 GB one about 192 GB. MacStories hit that wall even on 256 GB: two large Flash-Next builds didn't fit until parts of them were moved to the SSD. Check a model against the real limit before you pick a memory size: the calculator has a Mac tab where you choose Mac Studio, the chip and the memory, and it grades every model against that share. The Apple Silicon section of the VRAM guide covers raising that limit.

Read the fastest numbers carefully

Check whether speculation was on

The most quoted M5 Ultra figure is 108 tokens a second on Qwen3.8 Flash-Next. MacStories ran that with multi-token prediction (MTP) at depth 3, which lets the model draft several tokens and check them in one pass. It's a real result, and most owners will want MTP on. Just compare it with other MTP runs, because a plain run on another machine is measuring something else. An owner who tested Flash-Next on a 256 GB M5 Ultra put MTP at "roughly 2x" the plain speed (r/LocalLLaMA). Our speculative decoding guide explains when that gain holds up.

Builds move results too

The M3 Ultra rows in the llama.cpp thread were measured on a 2023 build of llama.cpp, and the M5 rows on a 2026 one. Software alone made a difference there: on an M2 Ultra, the newer build with flash attention on raised 8-bit writing speed from 66.64 to 82.83 tokens a second. So the gap between the M5 Ultra and the M3 Ultra in that table is partly the chip and partly two years of llama.cpp. When you compare two Macs, compare runs on the same build, or treat the difference as rough.

Frequently asked questions

Is the M5 Ultra twice as fast as the M5 Max for local LLMs?

No. It has twice the memory bandwidth, but in llama.cpp's standard test it writes text about 1.64 times as fast, 119 against 72 tokens a second on an 8-bit model. It reads long prompts much faster, which is where the Ultra pays off.

Should I buy the 96 GB M5 Ultra or the 128 GB M5 Max?

The 96 GB Ultra has the faster memory but the 64-core GPU and less room, about 72 GB for a model by default. The 128 GB Max gives a model about 96 GB at half the bandwidth. If your models fit in 72 GB, the Ultra is faster. If they don't, the Max runs them and the Ultra can't.

How much memory can a model use on a 256 GB Mac Studio?

About 192 GB by default, three quarters of the machine. You can raise the limit with the iogpu.wired_limit_mb setting, at the cost of memory for macOS and your other apps.

Is an RTX 5090 faster than the M5 Ultra?

For a model that fits in its 32 GB, yes. MacStories measured Qwen3.8 27B writing at 59 tokens a second on a 5090 and 48 on the M5 Ultra. Once a model needs more than 32 GB, the 5090 has to use system memory and slows down sharply, while the Mac keeps going.

Why do some M5 Ultra speeds look so much higher than others?

Usually because speculation was on. Multi-token prediction and draft models can roughly double the writing speed, and many headline figures use them. Check whether a result says it was a plain run before you compare it with another machine.

Do I need the 512 GB Mac Studio?

Only for the very largest models. A 256 GB model already gives about 192 GB to the GPU by default. The 512 GB version arrives later than the others, and one 256 GB owner already finds the GPU slow for that much memory, so twice the memory buys room for bigger models at the same writing speed.