llamaperf

DGX Spark

NVIDIA · 128GB unified memory · 25 reports

See what fits on this GPU →

Use the calculator to check which models fit in 128 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
Tone: positive
reported speed:
43.9 tokens/s generation
quant:
NVFP4 (NVFP4)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-usevisionmultilingual

User reports Qwen3.8-Flash-Next (180B MoE, 7.31B active) at 43.9 t/s on a single DGX Spark, roughly four times the 11.2 t/s of Qwen3.8-27B dense on the same machine. Setup is vLLM with NVFP4 weights, FP8 KV cache, MTP depth 3, greedy sampling, thinking mode off, tensor parallel size 1, and 320 tokens written per test. The 123.53 GiB checkpoint fits in 121.7 GiB unified memory because the 51.2B-parameter PLE n-gram table (47.68 GiB) is read from NVMe per token while the 120.8B experts (56.25 GiB at NVFP4) stay resident. Tool calls, thinking mode, and image input work on the live server; video input was accepted but the test clip was faulty so that result is inconclusive.

Sep 12, 2026

Ling-3.0 124B Flash

DGX Spark · llama.cpp · 262,144 ctx

reported speed:
35.7 tokens/s generation
quant:
Q5_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-context

User compares generation throughput on a 128 GB DGX Spark at three starting-context depths for two deployment paths. Official INT4 on a vLLM fork reaches 38.3 t/s short, 7.9 t/s at ~45K and 4.6 t/s at ~90K. Q5_K_M on llama.cpp reaches 35.7 t/s short, 33.6 t/s at ~45K and 33.2 t/s at ~90K. The config is 262,144 tokens of max context with ~103 GB total memory used, and streaming separates TTFT from decode. These are the creator's measurements, not an independent rerun, and two variables change at once (runtime and quantization), so no pure engine effect is isolated. The later 131,072-context serving default should not be attached to this table.

Sep 11, 2026
reported speed:
40.9 tokens/s generation
quant:
INT4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User benchmarks Ling-3.0-flash on a DGX Spark at 40.9 t/s on a short coding task. Setup is CUDA graphs with MTP n=1. Prose throughput is 38.7 t/s at 512-token output and 37.3 t/s at 2048-token output. MTP n=2 and n=3 are slower for prose. Acceptance lengths are 1.87 at n=1, 2.39 at n=2 and 2.77 at n=3. Without MTP the baseline is 22.9 t/s with CUDA graphs and 20.8 t/s eager.

Sep 7, 2026
Tone: positive
reported speed:
49.4 tokens/s generation · 2050.0 tokens/s prompt processing
quant:
MXFP8 x MXFP4
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash on two ASUS GX10 (DGX Spark) units connected via RoCE with TP=2, decoding 49.4 t/s at 4K context, 43.0 t/s at 16K, 37.9 t/s at 32K, 42.5 t/s at 128K and 39.8 t/s at 256K, with prefill of 2050 t/s, 2150 t/s, 2130 t/s, 1920 t/s and 1680 t/s at the same context lengths. Setup is a vLLM fork from local-inference-lab with MXFP8 x MXFP4 weights. At concurrency 4 and 128K context the aggregate is about 40-42 t/s. The KV cache holds up to about 1M tokens of context, though the user typically runs 256K. The user reports roughly 280W at maximum load and says the model beats M2.7 and Stepfun 3.7 on a private benchmark, praising its high-context retrieval and reasoning. A Docker compose file is provided.

Sep 7, 2026
Tone: positive
reported speed:
40.0 tokens/s generation
quant:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash at about 40 t/s single-stream and 350 t/s aggregate with 32 concurrent requests at 256k context on dual DGX Sparks. The 350 t/s figure is an aggregate across the 32 concurrent requests. User compares this to an RTX Pro 6000 at about 46 t/s and a Mac M2 Ultra 192GB at about 29 t/s, both at Q2.

Sep 7, 2026
Tone: positive
reported speed:
40.0 tokens/s generation · 1800.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports DeepSeek V4 Flash at 1800 t/s prefill and 40 t/s generation on two DGX Spark units. User praises the scalability via ConnectX and the power efficiency.

Sep 7, 2026

GLM-5.2

DGX Spark · vLLM · 131,072 ctx

Tone: positive
reported speed:
14.8 tokens/s generation · 512.0 tokens/s prompt processing
quant:
NVFP4
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextcoding

User reports GLM-5.2 NVFP4 at about 15 t/s decode on short context and about 13 t/s at long context on a 4x DGX Spark setup, with about 512 t/s prefill. Setup is 128K context with TP4/PP1/DCP4/MTP1 and an fp8 KV cache.

Sep 7, 2026
Tone: positive
reported speed:
20.6 tokens/s generation · 534.9 tokens/s prompt processing
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports REAP-pruned DeepSeek-V4-Flash served at 262k context on a single DGX Spark (Ascent GX10). Setup is vLLM. Prefill and generation throughput remain stable from 4K to 162K context. Benchmarks include pp and tg at various context lengths and concurrency levels.

Sep 7, 2026
Tone: positive
reported speed:
19.1 tokens/s generation · 1000.0 tokens/s prompt processing
quant:
2-bit (FP8)
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731 at 155.43 GiB FP8 across 48 shards on a DGX Spark (GB10, aarch64, sm_121, 121.7 GiB unified memory, 128 GiB swapfile), with prefill steady at 1000 tps and decode at about 19 tok/s. Setup is vLLM-Moet (vLLM v0.25.0 plus patch) with 2-bit MoE experts, a KV cache of 4.56M tokens at 512K context and util 0.90, and the MTP head from ycui7/DeepSeek-V4-Flash-MTP. With the MTP head decode reaches 26.6 tok/s, 48% over plain. Aggregate MTP versus plain is 25.2 versus 19.1 at conc1, 30.9 versus 26.4 at conc2, and 43.2 versus 45.6 at conc4. Per-request MTP versus plain is 25.2 versus 19.1 at conc1, 23.0 versus 17.7 at conc2, and 14.5 versus 13.9 at conc4. The 2-bit planes are 43 layers x 1.69 GiB, about 73 GiB. Boot takes about 10 min warm and 31-46 min cold.

Sep 7, 2026
Tone: positive
reported speed:
90.0 tokens/s generation · 1000.0 tokens/s prompt processing
quant:
MXFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports the MXFP4 version from Bartoswski on a DGX Spark, with prefill at 1K t/s and generation averaging 90 t/s. User calls it very efficient and the best scoring yet, and notes it lacks vision.

Sep 7, 2026
Tone: positive
reported speed:
24.0 tokens/s generation
quant:
UD-IQ3_XXS
kv:
bf16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports DeepSeek-V4-Flash-0731 UD-IQ3_XXS at roughly 23-27 t/s on a DGX Spark GB10, stabilizing at about 24 t/s. Setup is llama.cpp with DSpark speculative decoding. The user notes it is faster than Qwen3.6-27B-NVFP4 for codegen.

Sep 7, 2026
Tone: positive
quant:
NVFP4

User reports a C++20 port of the vLLM serving stack on a DGX Spark with Qwen3.6-27B NVFP4, reaching 86.05 to 1095.01 output tokens/sec across concurrency 1 to 32. The port runs slightly ahead of vLLM in the same tests. The user also reports DeepSeek-V4-Flash in 2-bit GGUF at 18.69 tok/s on the Spark, and compares against llama.cpp and MLX-LM.

Sep 7, 2026
reported speed:
82.0 tokens/s generation · 1400.0 tokens/s prompt processing
quant:
FP8
kv:
fp8_ds_mla

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731 (304B MoE) at 82 tok/s decode and about 1400 tok/s prefill on 2x DGX Spark, each with 128GB unified memory. Setup is vLLM 0.26.1rc1 with TP=2, 1M context and FP8. User asks for advice on reducing VRAM and OS RAM usage.

Sep 7, 2026
Tone: positive
reported speed:
60.0 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports 60 t/s on a 2x DGX Spark cluster with a vLLM recipe. The setup uses a 1M context window. User praises NVFP4 support and performance improvements, and compares the cluster favorably to Strix and M5 for prompt processing.

Sep 7, 2026

Muse 30B Glimmer

DGX Spark · llama.cpp · 1,000,000 ctx

Tone: positive
reported speed:
37.0 tokens/s generation · 390.0 tokens/s prompt processing
quant:
K-Quant-Dynamic (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagenticcodingvision

User reports retrieval at 832K tokens with 3/3 success. DFlash speculative decoding gives roughly 3x speedup. RPC split is slower than single-node.

Sep 7, 2026
Tone: positive
reported speed:
74.8 tokens/s generation · 1800.0 tokens/s prompt processing
quant:
NVFP4
kv:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports speculative decoding with DSpark, prefill of roughly 1.7-1.9k tok/s and a TTFT of 42s on a 64k prompt. KV capacity is 2.74M tokens. In an agentic eval through Codex CLI integration, the model landed a real feature in a 1,831-file codebase.

Sep 7, 2026
Tone: positive
reported speed:
40.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks DeepSeek V4 Flash at 16.5 t/s on a DGX Spark. Setup is llama.cpp with Q5_K_M, Q4_K_M and Q6_K quants. Q5_K_M is fastest at 40.2 t/s, Q4_K_M reaches 38.2 t/s and Q6_K reaches 32.0 t/s.

Sep 7, 2026
Tone: positive
reported speed:
38.7 tokens/s generation
quant:
INT4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 35.2 t/s with a community GGUF and a 2.4x speedup compared with DeepSeek V4 Flash. Official INT4 quants were initially reported as not running on a single Spark, later corrected to be the fastest path.

Sep 7, 2026
reported speed:
51.1 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 51.06 t/s single-stream and 184.03 t/s aggregate at c6 on 2x DGX Spark in TP2. Setup is SGLang with QSA, NEXTN and FlashInfer GDN. Disabling NEXTN gives 26.4 t/s at c1 and 111.8 t/s at c6. Medium thinking scored 22/24 on a 24-task LiveBench-style pilot.

Sep 7, 2026
reported speed:
49.7 tokens/s generation · 2875.0 tokens/s prompt processing
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 49.7 t/s decode on structured output and 34.8 t/s decode on prose with TP2 across two DGX Sparks, plus prefill of about 2,875 t/s for 11k tokens. Setup is MTP k=3 in eager mode. MTP acceptance is about 3.8 tokens/step on structured output. A single node reaches about 35 t/s on structured output.

Sep 7, 2026
Tone: positive
reported speed:
181.0 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports aggregate throughput of 195 t/s peak across about 9 concurrent agent sessions on a 2-node DGX Spark cluster, with single-stream decode at 30-50 t/s. The model uses hybrid attention (3/4 linear plus 1/4 sparse full), a 512-expert MoE, and MTP speculative decoding, with the PLE table mmap'd from NVMe. Context is stretched to 512K with YaRN.

Aug 29, 2026

GLM-5.2

DGX Spark · vLLM · 131,072 ctx

Tone: positive
reported speed:
24.0 tokens/s generation · 475.0 tokens/s prompt processing
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MTP3 as the default and MTP4 as the peak. Setup includes a bug fix for a draft parallel config missing a DCP copy. Prefill reaches about 475 tps and decode at bs=3 reaches about 48 tps.

Aug 28, 2026

Kimi K3

16× DGX Spark · vLLM

reported speed:
20.0 tokens/s generation · 750.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports the full Kimi K3 model running on a 16x DGX Spark cluster at an average of 20+ t/s, peaking at 38 t/s, with prefill at 750 t/s. This is a first run, and the user plans to optimize and publish a vLLM image.

Aug 28, 2026

User is setting up a 16x DGX Spark cluster to run frontier models locally. User mentions DeepSeek V4 pro, Kimi K3, GLM 5.5, and Minimax M4 as future models. No benchmark numbers are provided.

Aug 3, 2026
Tone: positive
quant:
NVFP4

User reports GLM-5.1-NVFP4 at 434 GB served across a 16x DGX Spark cluster with unified memory at TP=8. User plans to test DeepSeek and Kimi on the same cluster. User plans a future prefill/decode split using M5 Ultra Mac Studios.

May 1, 2026