llamaperf
← All articles

Hardware planning

RTX Spark for local LLMs: what DGX Spark tells you

By llamaperf · · 7 min read

Quick answer

RTX Spark puts the DGX Spark recipe, a Grace CPU, a Blackwell GPU and up to 128 GB of unified memory, into Windows laptops and small desktops. Expect DGX Spark speeds: around 59 tokens a second on a 120B mixture-of-experts model and under 30 on a dense 7B at 8 bits. NVIDIA hasn't published its memory bandwidth, and that number decides how fast it writes.

What NVIDIA announced

A DGX Spark you can carry

RTX Spark is NVIDIA's first chip for Windows PCs. It pairs a 20-core Grace CPU with a Blackwell GPU that has 6,144 CUDA cores, and it comes with up to 128 GB of memory that the CPU and GPU share. NVIDIA rates it at 1 petaflop of FP4 compute, the same figure it gives DGX Spark, and says laptops and compact desktops from ASUS, Dell, HP, Lenovo, Microsoft Surface and MSI arrive this fall (NVIDIA Newsroom). The headline claim for local models is this one:

"run 120-billion-parameter large language models with 1 million tokens context" (NVIDIA Newsroom)

That's a claim about fitting a model, and 128 GB does fit one. How fast it answers is a separate question.

The number NVIDIA left out

The release doesn't give a memory bandwidth, and for local models that's the number to want. Every token you get back means reading the model's active weights out of memory once more, so bandwidth sets a ceiling on writing speed. DGX Spark lists 273 GB/s on a 256-bit memory interface (NVIDIA DGX Spark). Laptop makers' spec sheets, as reported in a thread on NVIDIA's own forum, list faster LPDDR5X at 9,400 to 9,600 MT/s (NVIDIA developer forum). If the bus stays 256 bits wide, that works out to roughly 300 GB/s. That's our arithmetic, and NVIDIA hasn't confirmed the bus width, so treat it as a guess until someone measures one.

What DGX Spark already measures

Plain runs on the closest relative

The llama.cpp project keeps a set of llama-bench results for DGX Spark, run without speculation (llama.cpp DGX Spark results). They're the best guide to RTX Spark until owners post their own.

ModelSize on diskReads a 2,048-token prompt (t/s)Writes (t/s)
gpt-oss 20B, MXFP411.27 GiB4,505.8283.43
gpt-oss 120B, MXFP459.02 GiB2,443.9158.72
Qwen3-Coder 30B-A3B, Q8_030.25 GiB2,986.9761.06
Qwen2.5-Coder 7B, Q8_07.54 GiB2,250.2829.43

Look at the last two rows. The 30B model writes twice as fast as the 7B one, because it's a mixture of experts and reads only about 3B weights for each token. The 7B model is dense and reads all of its 7.54 GiB every time.

Why dense models feel slow on it

It's the same on every machine in this class. Take a dense 70B model at about 4.5 bits a weight, which is roughly 40 GB. At 273 GB/s the GPU can read it about seven times a second at most, so you'd expect fewer than seven tokens a second in practice (that's an arithmetic example). A mixture-of-experts model of the same total size reads a fraction of that per token and stays usable. People's own DGX Spark reports, with the model and quant of each run, are on the leaderboard under the 128 GB unified memory tier, next to Strix Halo.

How it lines up against Strix Halo and the Mac

Same memory, different speeds

RTX Spark is entering a crowded corner: machines with around 128 GB of shared memory. Here's how their bandwidth compares, using the figures our calculator runs on.

MachineMemoryBandwidth
DGX Spark128 GB273 GB/s
AMD Strix Halo128 GB256 GB/s
Mac Studio, M5 Maxup to 128 GB614 GB/s
Mac Studio, M5 Ultra96 to 512 GB1,229 GB/s

The Macs write much faster, with two to four and a half times the bandwidth. RTX Spark's case rests on other things. Reading a prompt leans on compute more than bandwidth, and DGX Spark already reads a 2,048-token prompt into a 120B model at over 2,400 tokens a second. It also runs CUDA, which most inference software supports first. The CPU is Arm, though, so check that your own tools have Windows on Arm builds before you buy. And it's the only one of these that comes in a Windows laptop.

Price is moving too

Memory costs have pushed these machines up. NVIDIA raised DGX Spark's price from $3,999 to $4,699 in February, citing memory supply (NVIDIA developer forum). NVIDIA hasn't announced RTX Spark prices, so compare against what the laptops sell for once they're listed, and size models for any of them in the calculator before you pay.

How to read the first RTX Spark numbers

Check what kind of run it was

The first speeds will come from launch reviews and early owners, and some will be much higher than DGX Spark's. Before you believe a gap, check whether speculation was on. Multi-token prediction and draft models can double writing speed, and plenty of Spark results posted so far had one of them switched on. Our speculative decoding guide explains when that holds up for your own use.

Then check how many requests ran at once. A server total adds up every user, so say a review quotes 200 tokens a second across eight streams: that's about 25 for each person (that's a made-up example). The guide to serving several users covers how to read those totals.

Laptops have one more variable

A laptop can throttle on battery or in a quiet power mode, and memory-bound work like writing tokens slows down with it. Look for reviews that say whether the machine was plugged in and which power mode it was in. Without that, you can't tell whether a slow result came from the chip or from the power plan.

Frequently asked questions

Is RTX Spark the same chip as DGX Spark?

It's the same idea in a PC: a Grace CPU, a Blackwell GPU and up to 128 GB of shared memory, rated at 1 petaflop of FP4 like DGX Spark. NVIDIA hasn't published RTX Spark's memory bandwidth, so whether it writes text at the same speed is still open.

How fast will RTX Spark run a 70B model?

A dense 70B model at about 4 bits is roughly 40 GB, and DGX Spark's 273 GB/s can read that about seven times a second at most. Expect single-digit tokens a second for dense 70B models. Mixture-of-experts models of similar size run several times faster.

Should I buy a Strix Halo laptop now or wait for RTX Spark?

Their memory bandwidth is close, 256 GB/s for Strix Halo against 273 GB/s for DGX Spark, so writing speed won't change much. RTX Spark's draw is CUDA and NVIDIA's software. If your tools need CUDA and have Windows on Arm builds, wait for the first measurements. If they run on llama.cpp or Vulkan, Strix Halo does the same job today.

Is RTX Spark faster than a Mac Studio for local LLMs?

Not for writing text. An M5 Max has more than twice the bandwidth and an M5 Ultra more than four times, and writing speed follows bandwidth. RTX Spark is competitive on reading prompts and runs CUDA software the Mac can't.

What is the memory bandwidth of RTX Spark?

NVIDIA hasn't said. DGX Spark has 273 GB/s. Laptop spec sheets list faster LPDDR5X, and if the bus is the same 256 bits that would be roughly 300 GB/s, but that's an estimate until someone measures a shipping machine.

Can RTX Spark run a 120B model?

Yes, with 128 GB of memory. DGX Spark runs gpt-oss 120B at about 59 tokens a second in llama.cpp, and that's a mixture-of-experts model, which is why it's fast. The smaller RTX Spark laptops with 24 to 32 GB can't hold it.