llamaperf

RTX 4090

NVIDIA · 24GB · 17 reports

See what fits on this GPU →

Use the calculator to check which models fit in 24 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
Tone: mixed
reported speed:
12.6 tokens/s generation
quant:
UD-Q3_K_M (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports generation at 12.0-13.1 t/s and prompt processing at 212-224 t/s for 8k-32k tokens. Setup uses --n-cpu-moe 39 to offload experts to CPU, and requires 128 GB RAM. Agentic tasks take 15-26 minutes. No quality benchmarks were run.

Sep 7, 2026
Tone: negative
reported speed:
1.5 tokens/s generation
quant:
UD-IQ4_NL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DSpark draft model performance of 1-2 t/s with DeepSeek-V4-Flash on llama-server, compared to 30-40 t/s with the MTP draft model. Setup is an RTX 4090 and an RTX 6000 Pro with 120 GB total VRAM and 32 GB RAM, running the unsloth DeepSeek-V4-Flash-0731 137 GB Q4 UD-IQ4-NL quant with the dspark-DeepSeek-V4-Flash-0731-Q8_0.gguf draft model, --spec-type draft-dspark, --spec-draft-n-max 3, --device-draft CUDA1, --tensor-split 2,3, --n-cpu-moe 13, --main-gpu 1, --split-mode layer, --ctx-size 65536, --flash-attn on, --batch-size 2048 and --ubatch-size 2048. User asks for advice on the correct setup.

Sep 7, 2026
Tone: mixed
reported speed:
3.0 tokens/s generation · 30.0 tokens/s prompt processing
quant:
IQ4_XS

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash 0731 at about 30 t/s prompt processing and about 3 t/s generation on CPU with an RTX 4090 and a Tesla P40. Setup is the IQ4_XS quant at roughly 127 GB with MTP enabled, and layer-by-layer assignment because tensor splitting is not supported. The user mentions an Unsloth 4bit K_XL quant at roughly 144 GB but could not run MTP with it due to memory constraints.

Sep 7, 2026

Qwen3.8 27B

RTX 4090 · vLLM · 262,144 ctx

Tone: positive
reported speed:
68.4 tokens/s generation
quant:
FP8
kv:
FP8 E4M3
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports decode speeds on an RTX 4090 with custom-tuned FP8 kernels: no MTP about 19.7 t/s, MTP=3 about 49.7 t/s, MTP=5 about 53.2 t/s, and with custom kernels about 68.4 t/s. Setup uses MTP=5 speculative decoding. The numbers come from real Cursor sessions, not fixed-prompt A/B.

Sep 7, 2026
Tone: positive
reported speed:
72.0 tokens/s generation
quant:
Q5_K_M

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares Q5 and Q6 quants of Qwen3.8 27B on a three.js 3D arena game. Q6 took 3h32m at 6.01 t/s with 84,183 tokens, while Q5 took 13m23s at 71.98 t/s with 57,803 tokens. Qwen3.6 Q5 was also tested at 84.79 t/s, and Q5 with medium reasoning reached 95.87 t/s. The user notes Q6 ran at temp 0.6 instead of 1.0.

Sep 7, 2026
Tone: positive
reported speed:
8.0 tokens/s generation
quant:
IXQ2/Q3

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash (v4-flash-0731) at about 8 t/s on a single RTX 4090 with 64 GB RAM. Setup is the Blaze engine with Unsloth's IXQ2/Q3 checkpoint, keeping heavily utilized experts in RAM with the CPU. Throughput drops to 5 t/s on an expert miss that causes a disk read. Prefill is a work in progress.

Sep 7, 2026

Qwen3.8 27B

RTX 4090 · llama.cpp · 194,048 ctx

Tone: positive
reported speed:
40.0 tokens/s generation · 1650.0 tokens/s prompt processing
quant:
Q4_K_M (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 194,048 context tokens at ~1650 prompt t/s and ~40 t/s generation without MTP, and 131,584 context tokens at ~1650 prompt t/s and ~65 t/s generation with MTP, up to 80 t/s for code, on an RTX 4090 FE with 48 GB DDR4 in WSL2. Setup uses a q8_0 KV cache and q4_0 draft KV cache, with --cache-ram 32768, --ctx-checkpoints 32, --no-mmap and --mlock.

Sep 7, 2026

Qwen3.8 27B

RTX 4090 · llama.cpp · 92,160 ctx

Tone: positive
reported speed:
134.0 tokens/s generation
quant:
Q4_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B on an RTX 4090, with throughput ranging from 44 t/s to 134 t/s and no RAM spilling. Setup includes an mmproj file for vision.

Sep 7, 2026
reported speed:
80.0 tokens/s generation · 1200.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares an RTX 4090 and an RTX 5090 for speed factors. User mentions Muse Glimmer as an alternative but has not benchmarked it.

Sep 7, 2026

Qwen3.8 27B

RTX 4090 · llama.cpp · 242,760 ctx

Tone: positive
reported speed:
90.0 tokens/s generation
quant:
IQ4_XS (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 90 t/s with MTP enabled and 45-50 t/s without MTP at full context. Setup uses a Q8 KV cache. User sees no difference versus FP16 and prefers speed over context.

Sep 7, 2026
Tone: positive
reported speed:
180.0 tokens/s generation · 5000.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports prefill speeds of ~5000 pp/tg for DeepSeek V4 Flash on 4x RTX 4090 with a vLLM fork. User also mentions Qwen3.8 Flash next with ~7500pp/135tg (MTP). User compares to llama.cpp and ik_llama, noting vLLM is much faster for prefill.

Sep 7, 2026
Tone: mixed
reported speed:
10.9 tokens/s generation · 132.5 tokens/s prompt processing
quant:
UD-Q2_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User also tests the IQ4_NL quant, reaching 50.7 t/s prompt and 8.1 t/s generation. User notes that with fixes such as flash attention, batch adjust and context quantization, the model could be decent on 24 GB GPUs. User compares it to Qwen 3.6 27B Q4_K_XL, which is faster but less reasoning, and prefers Qwen 3.6 27B for agentic tasks due to speed.

Aug 28, 2026

Qwen3.8 27B

RTX 4090 · 130,000 ctx

Tone: positive
quant:
Q4 (gguf)
codingcreative-writing

User reports running Qwen3.8-27b Q4 on an RTX 4090 to vibecode a Minecraft clone, with the model handling coding, audio, textures and 3D models. User is impressed with the capability and the cost, which came in under $1.

Aug 28, 2026
Tone: positive
reported speed:
138.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User benchmarks Gemma 4 26B-A4B against Gemma 4 12B on an RTX 4090. The 26B-A4B runs at 138 tok/s using 15 GB of VRAM, while the 12B runs at 80 tok/s using 9 GB. The 26B-A4B wins every scene and runs about 1.7x faster. The user notes the 12B is ideal for a 16 GB laptop.

Jun 4, 2026

Qwen3.6 27B Heretic-v2

RTX 4090 · llama.cpp · 262,000 ctx

Tone: positive
reported speed:
80.0 tokens/s generation
quant:
Q4_K_M (gguf)
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MTP draft acceptance of about 73%. Setup is a TBQ4_0 KV cache with an MTP draft of 3, using a fork of llama.cpp.

May 9, 2026
Tone: positive
reported speed:
149.6 tokens/s generation · 15.6 tokens/s prompt processing
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

text-generation

User reports about 150 t/s generation. The model is described as a star performer.

May 1, 2026

Qwen3.6 27B

RTX 4090 · LM Studio · 120,000 ctx

Tone: mixed
reported speed:
25.0 tokens/s generation
quant:
Q4_0 (gguf)
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports 3000 tokens in about 2 minutes, or 25 t/s, and is seeking faster performance. Setup is Q4_0 with 120k context and both caches quantized to 4_0. A reply adds a vLLM benchmark on an RTX 3090: a 27B INT4 quant at 125K context with a TurboQuant 3-bit NC KV cache and MTP speculative decoding, reaching 82 tok/s generation and 0.3-0.6s TTFT.

May 1, 2026