llamaperf

NVIDIA CMP 170HX 40GB (unlocked)

NVIDIA · 40GB · 3 reports

Engines people use on it: llama.cpp 1 · vLLM 1

Run models on your NVIDIA CMP 170HX 40GB (unlocked)? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA CMP 170HX 40GB (unlocked)

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 40 GB of VRAM.

reported speed:
96.9 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagentic

User reports Qwen3.8-Flash-Next at 96.9 tok/s single-request decode on four NVIDIA CMP 170HX 40GB cards. Setup is llama.cpp with UD-Q4_K_XL weights, q8_0 KV cache, MTP speculative decoding (draft length 4, GPU sampling), 262K context, and the SM clock pinned at 1410 MHz. The same configuration reaches 87.7 tok/s at 70K context; three concurrent 3K requests give 34-37 tok/s each (about 100 tok/s aggregate). Earlier revisions of the fork measured 64-73 tok/s at 2K and 58-70 tok/s at 70K, against a baseline fork at 46-55 tok/s and 27-45 tok/s respectively.

Sep 27, 2026
Tone: positive
reported speed:
225.0 tokens/s generation
quant:
W4A16
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at around 225 t/s on a single CMP 170HX with 40GB VRAM at 132k context. Setup is vLLM with W4A16 4-bit weights and an fp8 KV cache. The GPU is overclocked to NDIV 60, giving 1.89 TB/s memory bandwidth at around 1500 MHz, and has been stable for over a week at 65-67C. The user also runs Minimax H3 for video generation on the same card and notes NDIV 62 crashes while NDIV 60 is stable.

Sep 24, 2026
Tone: positive
reported speed:
202.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B token generation rising from 110 T/S to 202 T/S on a CMP 170HX 40GB after overclocking. Nothing changed except the overclock, which raised memory bandwidth from 1,386.2 GB/s to 1,890.1 GB/s, a +36.4% increase. GPU wattage is 300 watts and GPU temps are slightly lower after the overclock.

Sep 20, 2026

Get a weekly email of new NVIDIA CMP 170HX 40GB (unlocked) reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 125B (6B active) Flash-Next
4× NVIDIA CMP 170HX 40GB (unlocked)
UD-Q4_K_XL
llama.cpp
262,14496.9 tokens/s
Qwen3.8 27B
NVIDIA CMP 170HX 40GB (unlocked)
W4A16
vLLM
132,000225.0 tokens/s
Qwen3.8 27B
NVIDIA CMP 170HX 40GB (unlocked)
Not reported
Engine not reported
Not reported202.0 tokens/s