llamaperf

RTX Pro 6000 Blackwell

NVIDIA · 96GB · 27 reports

See what fits on this GPU →

Use the calculator to check which models fit in 96 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

User announces the release of Aurora1.0-150M, a 150M parameter model trained on 7B tokens using an RTX Pro 6000 Blackwell. No inference engine, quantization, context length, or throughput figures are reported. Benchmarks cited: PIQA 62.24%, Hellaswag 32.20%, Arc-Easy 44.91%, Arc-Challenge 25.00%, Arithmark 3.0 33.90%, CapitalBench 36.55%.

Sep 13, 2026
reported speed:
20.3 tokens/s generation
quant:
IQ4_XS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User compares llama.cpp, SGLang, and FreeToken on Qwen3.8-Flash-Next. At full context, TTFT is 35.4s for SGLang, 80.4s for FreeToken, 210.2s for llama.cpp with MTP, and 258.4s for llama.cpp baseline. Decode at full context is 126.9 t/s for SGLang, 87.5 t/s for FreeToken, 52.6 t/s for llama.cpp with MTP, and 20.3 t/s for llama.cpp baseline. MTP improves decode 1.63x at 8K and 1.69x at 32K. Accuracy is 95.22-95.75% on GSM8K and 92.20-93.00% on MATH-500. Startup takes 16s for llama.cpp, 108s for SGLang, and 126s for FreeToken.

Sep 8, 2026
reported speed:
57.0 tokens/s generation
quant:
BF16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Muse Glimmer 30B at ~57 t/s with speculative decoding using DFlash, against ~25 t/s without speculation, a 2.3x speedup. Setup is vLLM with six patches applied to the image. Mean acceptance length is ~2.5 tokens per verification step, with overall draft acceptance ~10%. Per-position acceptance is ~73% at position 0, ~40% at 1, ~15% at 2, and near zero past 5. The recipe's 3.1x figure was greedy decoding with K-quant on llama.cpp under different conditions. The predicted ceiling is ~65 t/s at 2.6 mean acceptance.

Sep 7, 2026
reported speed:
35.0 tokens/s generation
quant:
Q8_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks TensorSharp against llama.cpp on an RTX PRO 6000 Blackwell, with plain text generation decode at about 35 tok/s at 60 prompt tokens. The user also tests DFlash speculative decoding and 2x RTX PRO 4000 Blackwell 24GB with tensor parallelism.

Sep 7, 2026
Tone: positive
reported speed:
29.0 tokens/s generation · 1328.0 tokens/s prompt processing
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports prefill speeds of 152 t/s at ~1K, 321 t/s at 2,043, 906 t/s at 8,623, 1,328 t/s at 23,348 and 1,204 t/s at 62,403 with INT4 experts. Decode after a ~1K prompt runs 29.4, 28.2 and 28.5 t/s for 50, 100 and 250 tokens; after a 62K prompt it is 19.4 t/s. The run keeps 6,440 of 11,008 routed experts resident.

Sep 7, 2026
Tone: positive
quant:
BF16
rating:
4/5

User reports Qwen 3.6 27B abliterated BF16 running locally via vLLM with llama-swap and MTP speculative decoding, scoring 8/10 on a Terminal-Bench 2.0 pilot. The same setup beats DeepSeek-V4 IQ2 at 7/10 and comes close to an FP8 API at 9/10. It is the only configuration to pass the cancel-async-tasks hard task, and it missed build-cython-ext and sqlite-db-truncate on timeout.

Sep 7, 2026
Tone: mixed
reported speed:
45.0 tokens/s generation
quant:
Q4_K_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731-UD-Q4_K_XL at 45 t/s and 512k context as their primary coding model, across a multi-GPU setup of an RTX 6000 Pro Blackwell 96GB, 2x RTX 5090, an RTX 4090, 3x AMD R9700, and a Strix Halo laptop. Other models in use are GLM-4.7-flash, Gemma-4-12B-it, KAT-Coder-V2.5-Dev, LFM2.5-VL-1.6B, GLM-5.2-UD, Kimi-K2.7-Code, and MiniMax-M3. The user is dissatisfied with the workflow, describing it as a good fast worker or a good thinker but not both, and notes that switching models kills the cache.

Sep 7, 2026
Tone: positive
reported speed:
31.2 tokens/s generation
quant:
Q4_K_XL
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports DSpark speculative decoding with the drafter in RAM running faster than in VRAM. Setup uses a q8_0 KV cache, which allows 768K context with minimal decode loss. LiveCodeBench scores 28/30 (93.3%).

Sep 7, 2026

Qwen3.8 27B

RTX Pro 6000 Blackwell · vLLM · 262,144 ctx

Tone: mixed
reported speed:
114.8 tokens/s generation
quant:
FP8
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares Qwen3.8-27B-FP8 against Qwen3.6-27B-FP8 on an RTX PRO 6000 Blackwell with vLLM. In an MTP sweep, Qwen3.8 peaks at MTP 6 with 114.8 tok/s and Qwen3.6 peaks at MTP 7 with 131.6 tok/s. Qwen3.8 is 5-20% slower across MTP steps, though quality is comparable or better at some steps. A pre-sweep manual run at MTP 5 showed Qwen3.8 faster, at 108.9 vs 103.6 tok/s.

Sep 7, 2026
Tone: positive
reported speed:
99.6 tokens/s generation
quant:
Q8_0 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports that tuning the MTP draft window from n-max 12 to 5 improved throughput from 63 to 99.58 t/s. The Q4_K_XL quant reached 125.63 t/s, faster than Q8_0 at 99.58 t/s, with no quality gap detected. The user recommends reasoning_effort medium over xhigh or low, and notes that model quality revealed harness bugs.

Sep 7, 2026

Qwen3.8 27B

RTX Pro 6000 Blackwell · NInfer · 262,144 ctx

Tone: positive
reported speed:
96.8 tokens/s generation
quant:
groupwise-int
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User reports a long coding-agent session with DeepSeek Harness at a weighted 96.8 t/s streaming rate and a median of 104.4 t/s. The model is a groupwise-int artifact, not NVFP4. Context compaction worked well.

Sep 7, 2026

Qwen3.8 27B

RTX Pro 6000 Blackwell · NInfer · 262,000 ctx

Tone: positive
reported speed:
104.8 tokens/s generation · 12400.0 tokens/s prompt processing
quant:
groupwise-int

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentictool-uselong-context

User reports Qwen3.8-27B on an RTX PRO 6000 Blackwell at 104.83 tok/s weighted decode, running a long-horizon agentic task at 262K context. Setup is NInfer with a groupwise-int quant (mixed Q4/Q5/Q6). The run covered 966 model calls, 131.2M input tokens and 853.3K output tokens. Prompt processing was about 12.4K tok/s, measured client-side and including queue time. The user plans to try an NVFP4 profile and estimates API cost savings of $650+.

Sep 7, 2026
Tone: positive
quant:
Q8

User benchmarks Qwen3.6-27B at 32 concurrent clients on the Paddock engine, with TTFT of 697 ms against 2.5 s for vLLM and 6.9 s for llama.cpp. User also reports Qwen3.8-27B at Q8 on an RTX PRO 6000 with speculation on and off: 46.8 to 202 tok/s single stream, 320 to 822 at eight concurrent chats, and 1005 to 1285 at 32. The engine is free but not open source.

Sep 7, 2026

Qwen3.8 27B

RTX Pro 6000 Blackwell · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
153.9 tokens/s generation
quant:
Q4_K_M (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks DFlash 2 speculative decoding on Qwen 3.8 27B with a llama.cpp PR build on an RTX PRO 6000 Blackwell 96GB and a Ryzen 9 9950X. Setup is llama.cpp with 262144 context, an f16 KV cache and full offload. The user reports 2.26x speedup on LiveCodeBench, from 67.97 to 153.91 tok/s, and 4.68x on multi-turn coding with n-gram lookup, and gives a detailed analysis of n-gram drafters and draft width.

Sep 7, 2026
reported speed:
114.2 tokens/s generation
quant:
Q6_K (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User compares Qwen 3.8 27B community quants on an RTX Pro 6000 96GB, with atomic ad-q6_k averaging 114.17 t/s as the best. The quants tested are atomic ad-q6_k, unsloth ud-q6_k_l and bartowski q6_k. The user also compares the results against a Claude Opus 4.6 subscription, with prompts generating self-playing 3D HTML games.

Sep 7, 2026
Tone: positive
reported speed:
110.0 tokens/s generation
quant:
NVFP4 (safetensors)
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingtool-usevisionagentic

User reports Qwen3.8-Flash-Next at roughly 110 t/s, ranging from 76 to 125 t/s, on a single RTX Pro 6000 Blackwell with 96 GB, at 165,000 to 170,000 tokens of context. Setup is vLLM with the NVFP4 quant and INT4 quantized PLE n-grams, MTP speculative decoding, prefix caching, and PLE CPU offload. The user praises quality for coding and agentic tasks.

Sep 7, 2026
Tone: mixed
reported speed:
177.0 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares Flash-Next against a dense 27B model. Flash-Next is faster and mechanically flawless, with zero failures on strict JSON, injection resistance and SLAs, and wins the high-reasoning spatial and code-gen tier with better failure modes. The dense 27B still wins sustained multi-step symbolic work such as bug-fixing, math proofs and abstract puzzles. On that symbolic work Flash-Next shows a new failure shape: it promises the deliverable, declares "done", and outputs nothing. Both models use the same reasoning_effort knob with radically different semantics, so Flash-Next is not a drop-in replacement but a conditional promotion.

Sep 7, 2026
reported speed:
109.1 tokens/s generation · 1955.0 tokens/s prompt processing
quant:
UD-IQ4_XS (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8 Flash Next across VRAM tiers from 8 GB to 96 GB with llama.cpp b10666. CPU-only decode reaches 8.34 tok/s and full 96 GB decode reaches 109.07 tok/s. At 245K context, decode ranges from 14.89 tok/s on 24 GB to 21.61 tok/s on 96 GB. The model activates 6B params per token (MoE). RAM-resident loading gives 1.87x prefill versus mmap. Non-unified KV reaches 92.0 tok/s total at concurrency 16. A PLE table on CUDA causes a 55.6x decode slowdown, from 108.5 to 1.95 tok/s.

Sep 7, 2026
Tone: positive
quant:
FP8

User benchmarks Qwen3.8-27B FP8 on one RTX PRO 6000 in Paddock, an open-source inference engine in Rust/C++ with its own CUDA kernels, reaching 1062 tok/s at 32 clients against 958 for vLLM and 844 for SGLang. Paddock is faster than vLLM in 13/13 cells, by 1.02x to 1.19x, and faster than llama.cpp Q8_0 in 13/13, by 1.5x to 37x.

Sep 7, 2026
Tone: positive
quant:
FP8

User benchmarks Qwen3.8-27B FP8 on one RTX PRO 6000 in a new open-source Rust/C++ inference engine with its own CUDA kernels. The engine supports GGUF and safetensors, is CUDA only, and is validated on Blackwell and Ampere. The user claims it is faster than vLLM in 13/13 cells (1.02x-1.19x), faster than SGLang in 10/13, and faster than llama.cpp Q8_0 in 13/13 (1.5x-37x). At 32 clients with 1024 in/1024 out it reaches 1062 tok/s against vLLM at 958 and SGLang at 844.

Sep 4, 2026
Tone: positive
agenticcodinglong-context

User reports Mimo 2.5 stays fast at large context on dual RTX Pro 6000, using 5-to-1 sliding-window attention similar to Gemma 3. User also mentions Step 3.7 Flash with 3-to-1 sliding-window attention at ~40 t/s at 178k context. User notes MiniMax M3 and DeepSeek V4 have kernel issues on Blackwell consumer GPUs.

Aug 28, 2026
reported speed:
57.4 tokens/s generation
quant:
Q6_K (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks several quants of a model on 2x RTX PRO 6000 Max-Q: Q4_K_M at 67.3 t/s, Q3_K_L at 63.3 t/s, and IQ2_M at 78.7 t/s. User also benchmarks Nemotron-Labs-Audex-30B-A3B, a 30B MoE with ~3B active, on the same hardware: Q8_0 at 287 t/s, Q5_K_M at 334 t/s, Q4_K_M at 345 t/s, and MXFP4_MOE at 329 t/s.

Aug 28, 2026
Tone: positive
reported speed:
67.0 tokens/s generation
quant:
AD-Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

creative-writing

User compares AD-Q4_K_M, AD-Q5_K_M, AD-Q6_K, and Q8_0 quants of Qwen 3.8 27B on an RTX PRO 6000, reporting top-1 accuracy and KLD against BF16. Generation speeds are 67, 57, 49, and 50 tok/s respectively. The user recommends AD-Q6_K as the safest pick.

Aug 28, 2026

User benchmarks Qwen3.6 27B BF16 and Qwen3.6 35B BF16. For the 35B, best generation throughput is 3500 t/s at 128 concurrency with MTP off, and prompt throughput is 30000 t/s. The 27B was also tested with MTP on and off.

May 25, 2026
Tone: positive

User reports FastDMS, a DMS KV-cache compression implementation, decoding 1.5-2x faster than vLLM BF16/FP8 with 5-8x less KV memory. Benchmarks cover Llama-3.2-1B and Qwen3-8B DMS checkpoints, with KLD and token match comparable or better than vLLM's FP8/TurboQuant. Training took about 20 min on an RTX Pro 6000 Blackwell.

May 5, 2026
Tone: mixed
codingsummarization

User compares Qwen3.6-27B with and without thinking against Coder-Next. Qwen3.6-27B with thinking disabled is the most consistent at a 95.8% ship rate, and Qwen3.6-27B and Coder-Next are statistically tied overall. Qwen3.6-35B-A3B performed poorly and was dropped.

May 4, 2026