llamaperf

NVIDIA CMP 170HX 64GB (unlocked)

NVIDIA · 64GB · 10 reports

As of 7 Oct 2026, the models most run on the NVIDIA CMP 170HX 64GB (unlocked), with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: vLLM 7 · llama.cpp 2

Run models on your NVIDIA CMP 170HX 64GB (unlocked)? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA CMP 170HX 64GB (unlocked)

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 64 GB of VRAM.

reported speed:
117.0 tokens/s generation · 6066.0 tokens/s prompt processing
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4.1-Flash at 117 tok/s decode on one stream at 128k context and 6,066 tok/s prefill at 105k tokens on eight 64 GB CMP 170HX mining cards. Setup is vLLM with fp8 KV cache, PP=8, DSpark speculative decoding with 5 draft tokens, Engram tables in pinned host RAM, and 1M context. Decode drops to 96 tok/s at 512k. Eight concurrent streams reach 532 tok/s aggregate (66 per stream) at 105k. KV pool holds 6.17M tokens. One card was capped to 180 W after PCIe bus drops.

Oct 5, 2026
Tone: positive
reported speed:
101.3 tokens/s generation · 3191.0 tokens/s prompt processing
quant:
W4A16 (W4A16)
kv:
BF16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-contexttool-usemultilingual

User reports Qwen3.8-Flash-Next at 101.3 t/s decode on Japanese prose and 170.0 t/s on code, single stream, on two NVIDIA CMP 170HX cards. Setup is vLLM with W4A16 weights, BF16 KV cache, MTP k=4 speculative decoding, 262,144-token context, and expert parallel across the two cards. The PLE n-gram table was converted to FP8 locally to fit in 92 GiB of host RAM. Aggregate throughput reaches 392 t/s at 4 concurrent requests. Prefill measures 3,191 t/s at 6,954 tokens and 3,261 t/s at 27,853 tokens. A 200,087-token prompt returned the planted code with 70.5 s time to first token and 69.0 t/s decode at depth. MTP is worth about 1.6x on prose and 2.6x on code. The user notes xhigh reasoning effort spent the whole budget without answering in 5 of 6 runs.

Sep 27, 2026
reported speed:
95.0 tokens/s generation
quant:
FP8
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User reports DeepSeek V4 Flash Vision at up to ~95 tok/s on 4× CMP 170HX 64GB (256GB aggregate HBM2e). Setup is vLLM with pipeline parallel 4, max_model_len 262144, FP8 KV cache, DSpark speculative decoding with 6 tokens, and max_num_seqs 2. Normal coding workload runs ~50–70 tok/s, strong speculative-decoding periods ~80–90 tok/s, peak ~95 tok/s. Cards sit around 56–62GB VRAM each at 50–60°C. Next tests planned for Qwen3.8 Flash Next and GLM 5.3.

Sep 24, 2026
reported speed:
116.0 tokens/s generation · 6800-7950 tokens/s prompt processing
quant:
W4A16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next-80B at 116 t/s single-stream decode on a single NVIDIA CMP 170HX with 64 GB HBM2e, at a 150 W cap. Setup is vLLM 0.27.1 with W4A16 weights (40.9 GB), torch 2.13.0+cu130, CUDA 13.0, Ubuntu 26.04 LTS, on an AMD Ryzen Threadripper PRO 3945WX with 128 GB DDR4 ECC. Prefill over about 8.9k tokens measured 6800 to 7950 t/s; power draw 137 to 145 W. Aggregate throughput at 8 concurrent requests was 352 t/s, saturating at 4 slots. Also measured on the same card for context: Ornith-1.5-35B FP8 at 122.5 t/s and Qwen3.8-27B W4A16 with DFlash2 at 127 t/s single-stream.

Sep 23, 2026
reported speed:
98.1 tokens/s generation · 5300.0 tokens/s prompt processing
quant:
MXFP4+FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731 at 98.1 t/s single-stream decode on 4x CMP 170HX (64 GB HBM2e each), up from a 50.8 t/s baseline with DSpark speculative decoding. Setup is vLLM with pipeline parallelism, native MXFP4+FP8 weights (~155.4 GiB), context verified to 1,047,736 tokens. At 100k context single-stream decode is 38.8 t/s with 14.6 s time to first token. Aggregate throughput across 64 concurrent requests is 712.8 t/s with DSpark (472.0 t/s baseline), and 90.0 t/s at 100k context. Prefill across the 25k-77k context range is about 5,300 t/s.

Sep 23, 2026
Tone: positive
reported speed:
147.0 tokens/s generation
quant:
W4A16 (W4A16)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-context

User reports Qwen3.8-27B at 147.0 t/s single-stream decode on one NVIDIA CMP 170HX 64GB, with 134.7 t/s at 4K, 100.1 t/s at 65K, ~90 t/s at 126K, and 64.9 t/s at 250K context. Setup is vLLM 0.27.1 with a W4A16 target, a DFlash2 W4A16 drafter, FP8 target KV, BF16 draft KV, a custom SM80 split-KV verifier, full CUDA Graph, 35 verifier segments / 140 CTAs, k=3 draft tokens, MAX_SEQS=1, and 1350 MHz / 180W. These are decode-only numbers, not end-to-end throughput including prefill. FP8 beat INT8 at long context (48.3 vs 115.4 ms/iter at 250K). Allowing 2 concurrent long requests made makespan and slowest-request throughput worse at both 126K and 250K.

Sep 20, 2026
Tone: mixed
reported speed:
50.0 tokens/s generation
quant:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at around 50 t/s on a CMP 170HX, running at Q8 or BF16. The user finds the model hallucinates classes and APIs in 90%+ of answers to domain-specific enterprise Java questions, while DeepSeek and Microsoft Copilot produce working code 99% of the time. The user asks whether Qwen's agentic optimization is responsible and requests tips for a harness, system prompt, or skills to reduce hallucination.

Sep 18, 2026
Tone: positive
reported speed:
63.0 tokens/s generation · 1700.0 tokens/s prompt processing
quant:
Q6_K
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B at 63 t/s on 4x CMP 170HX 8GB cards flashed to 64GB each, 256GB total. Setup uses MTP with little-MoE default. Generation reaches 110 t/s with MTP optimistic.

Sep 7, 2026
reported speed:
29.0 tokens/s generation · 450.0 tokens/s prompt processing
quant:
Q4_K_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports testing multiple models on 4x CMP 170HX 64GB cards, including DeepSeek V4-Flash 0731 with a Q4_K_XL quant, 13B active MoE, and 1M context, plain no-spec. The cards are cut-down A100 mining cards with 64GB each, PCIe Gen2 x4, and no NVLink. Other models tested include gpt-oss-120B, Qwen3.6-35B-A3B, GLM-4.5-Air, and MiniMax-M2.7.

Sep 7, 2026
Tone: positive
reported speed:
3468.0 tokens/s prompt processing
quant:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports an unlocked CMP 170HX, A100 silicon with a firmware unlock, reaching 193 TFLOPS tensor throughput, up from 6.3 TFLOPS. In llama.cpp, pp512 rises from 599.6 to 3468 tok/s. Serving Qwen3.8-27B-FP8 under vLLM, one 170HX beats a 2x3090 tensor-parallel pair on prefill at 197W versus 454W.

Aug 28, 2026

Get a weekly email of new NVIDIA CMP 170HX 64GB (unlocked) reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
DeepSeek V4.1 Flash 552B (16B active)
8× NVIDIA CMP 170HX 64GB (unlocked)
Not reported
vLLM
1,048,576117.0 tokens/s
Qwen3.8 125B (6B active) Flash-Next
2× NVIDIA CMP 170HX 64GB (unlocked)
W4A16
vLLM
262,144101.3 tokens/s
DeepSeek V4 Flash Vision
4× NVIDIA CMP 170HX 64GB (unlocked)
FP8
vLLM
262,14495.0 tokens/s
Qwen3-Next 80B (3B active)
NVIDIA CMP 170HX 64GB (unlocked)
W4A16
vLLM
Not reported116.0 tokens/s
DeepSeek V4 Flash 284B (13B active)
4× NVIDIA CMP 170HX 64GB (unlocked)
MXFP4+FP8
vLLM
1,047,73698.1 tokens/s
Qwen3.8 27B
NVIDIA CMP 170HX 64GB (unlocked)
W4A16
vLLM
250,000147.0 tokens/s