llamaperf
← All articles

Hardware planning

Gorgon Halo for local LLMs: more memory, same speed

By llamaperf · · 5 min read

Quick answer

Gorgon Halo is Strix Halo with faster memory clocks and up to 192 GB instead of 128 GB. Peak bandwidth goes from 256 to 273 GB/s, about 7%, and that's the most a model that already runs on Strix Halo will speed up. The point of the chip is room. Up to 160 GB can go to the GPU, which is enough for mixture-of-experts models too big for Strix Halo.

What AMD changed

Same chip design, faster memory

The Ryzen AI Max 400 series, codenamed Gorgon Halo, keeps Strix Halo's layout: 16 Zen 5 cores, a 40-unit Radeon GPU and a 256-bit memory bus. On the flagship Max+ PRO 495 the GPU clocks 100 MHz higher and the memory is LPDDR5x-8533 (AMD Ryzen AI Max+ PRO 495). Strix Halo's Max+ 395 uses LPDDR5x-8000 on the same bus (AMD Ryzen AI Max+ 395).

The memory speed is the figure to look at for local models, because every token the model writes means reading its active weights out of memory again. Multiply it out from AMD's numbers and a 256-bit bus at 8,533 MT/s peaks at about 273 GB/s, against 256 GB/s at 8,000. That's our arithmetic, and it comes to about 7% more.

The real change is capacity

AMD sells the chip on memory:

"the Ryzen AI Max 400 Series offers up to 192GB, with up to 160GB dedicated to the GPU" (AMD blog, 28 September 2026)

Strix Halo stopped at 128 GB, of which the GPU could take up to 96 GB. Gorgon Halo adds 64 GB to both. AMD says its partners are selling systems now, and they cost real money: Framework's 192 GB desktop starts at $6,799 (Phoronix).

What speed to expect

Start from what Strix Halo measures

Nobody has posted a measured Gorgon Halo run yet. The closest evidence is Strix Halo itself, and kyuz0's public benchmark set runs the same models on it across every Vulkan and ROCm build (kyuz0 Strix Halo toolboxes). On Vulkan RADV with flash attention on, gpt-oss 120B reads a 512-token prompt at 719.91 tokens a second and writes at 56.61. Put 32,768 tokens in the context first and writing falls to 43.01.

Scale that writing speed up by the bandwidth gain and Gorgon Halo would manage about 60 tokens a second. That's arithmetic from the 56.61 figure, and you should read it as a ceiling, since memory rarely sustains its rated peak and a different driver or build could easily eat a 7% gain.

Dense models stay slow

gpt-oss looks quick because it's a mixture of experts, so each token reads only a small slice of its weights. A dense model reads every weight for every token. People report Qwen3.8 27B on Strix Halo at around 10 tokens a second, and a dense 27B on Gorgon Halo will land in the same place. These machines share one tier on the leaderboard, where you can see what people run on them and how fast, with the quant and engine behind each run. If it seems odd that total size and speed come apart like this, our guide to mixture-of-experts models explains why.

Who should buy one

Buy it for a model that doesn't fit today

The extra 64 GB earns its price when there's a specific model you want that Strix Halo can't hold. Large mixture-of-experts models are the obvious case. They need a lot of memory to load, yet each token reads only the active experts, so they stay usable at this bandwidth. Size the model you have in mind in the calculator before you spend the difference. It grades a model against each machine's memory and bandwidth.

If everything you run already fits in 96 GB, you'd be paying a lot for about 7% at best, and a Strix Halo box runs the same models at nearly the same speed. Judge Gorgon Halo as a memory upgrade, by the models it lets you load.

If you need more speed

AMD hasn't announced a part with a wider or faster memory bus, so this class of machine sits near 273 GB/s for now. If you mostly run dense models and want text back quickly, neither Strix Halo nor Gorgon Halo is the machine to buy. A Mac or a discrete GPU with faster memory will do better, and the comparison pages put a card side by side with these machines.

Frequently asked questions

Is Gorgon Halo faster than Strix Halo for local LLMs?

A little. Its memory runs at 8,533 MT/s instead of 8,000 on the same 256-bit bus, roughly 7% more bandwidth, so writing speed can go up by about that much at most. The bigger difference is memory: up to 192 GB against 128 GB.

How much memory can the GPU use on a Ryzen AI Max 400?

AMD says up to 160 GB, out of a possible 192 GB. On Strix Halo the GPU could have up to 96 GB of its 128 GB.

Should I buy a Strix Halo now or wait for Gorgon Halo?

There's nothing to wait for, since Gorgon Halo machines are on sale already. They're expensive, with Framework's 192 GB desktop starting around $6,800. If your models fit in 96 GB, a Strix Halo box runs them at nearly the same speed for less, so pay for Gorgon Halo when a particular model needs the extra memory.

Can Gorgon Halo run a 300B model?

Sometimes. At 4 bits a 300B model is about 150 GB of weights (an arithmetic example), so it fits under the 160 GB limit with little left over for context. It's only usable as a mixture of experts with few active weights per token. At 273 GB/s a dense model that size would write very slowly.

How fast is Gorgon Halo on gpt-oss 120B?

No one has measured it yet. Strix Halo writes at about 57 tokens a second in llama.cpp with Vulkan, and Gorgon Halo should come in a few percent above that, around 60 at most.