llamaperf

Qwen3.8

Alibaba · 50 reports

By engine

EngineAvg t/sRangeN
LM Studio43.035–553
llama.cpp41.50–154116
Ollama34.234–342
text-generation-webui32.533–331
ExLlamaV325.025–251
MLX16.94–295
vLLM121.450–38224
reported speed:
15.8 tokens/s generation · 81.8 tokens/s prompt processing
quant:
4bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark compares MLX vs llama.cpp on Qwen3.8-27B. MLX: mlx-community/Qwen3.8-27B-4bit, ~16.1GB, prompt 81.76 tok/s, generation 15.81 tok/s, peak memory 16.39GB. llama.cpp: unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M, 15.32 GiB / 27.32B params, full Metal offload, Flash Attention enabled, prompt 99.61 ± 0.44 tok/s, generation 9.69 ± 0.34 tok/s. llama.cpp ~22% faster prompt processing, MLX ~63% faster generation.

Qwen3.8 Flash-Next

M2 Max 96GB · oMLX · 32,768 ctx

Tone: negative
reported speed:
18.3 tokens/s generation · 235.7 tokens/s prompt processing
quant:
oQ4e (MLX)
kv:
8.0
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports slow performance with Qwen3.8-Flash-Next on M2 Max 96GB using oMLX engine. Benchmark results show pp1024/tg128: 124.1 pp TPS, 20.4 tg TPS; pp4096/tg128: 170.7 pp TPS, 16.7 tg TPS; pp8192/tg128: 214.3 pp TPS, 18.5 tg TPS; pp16384/tg128: 235.7 pp TPS, 18.3 tg TPS. Also tried oQ4e-fp16-mtp, oQ3-fp16-mtp, oQ3-MTP variants. fp16 variants even slower (~150pp, 8tg). User notes llama.cpp achieves 350-400 pp TPS. Model is MoE with 6B active parameters. Context length set to 32768. KV cache quant is turboquant_kv_bits 8.0. Engine is oMLX (custom MLX-based).

Qwen3.8 Flash-Next

M3 Max 48GB · oMLX · 65,536 ctx

Tone: mixed
reported speed:
38.0 tokens/s generation
quant:
T5 (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingmath

Runs on a custom oMLX fork with a Metal kernel for ternary experts; stock oMLX/mlx-lm won't load it. ~35.6 GiB resident: routed expert gate/up as ternary (Bonsai T5 packing, ~1.875 bpw), expert down projections as Q3, 53 GB n-gram table left on SSD as Q8 and mmapped per token. All low-bit tensors fitted with Unsloth's imatrix. Prefill 210-380 tok/s. 64K context confirmed (65,536-token prompt + 256 output at 30.7 tok/s, 42.3 GiB physical peak); 96K trips the prefill guard. Physical peak at 8K context ~41.5 GiB; swaps ~2 GiB once at load. Setup: oMLX memory guard 'safe' profile, limit 48 GB, one model, one request at a time. Quality vs Unsloth UD-Q4_K_XL (doesn't fit in 48 GB): KLD vs Q8_0 0.49 vs 0.036; MMLU 83.0% vs 89.7%; GSM8K 90.0% vs 92.0%; HumanEval 92.7% vs 95.7%. Known wart: sometimes ignores 'answer with just the letter' in Chinese.

Qwen3.8 27B GSQ-RCO

RTX 5060 Ti 16GB · beellama · 85,000 ctx

Tone: positive
reported speed:
45.0 tokens/s generation · 300.0 tokens/s prompt processing
quant:
IQ3_XXS (GGUF)
kv:
kvarn4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionlong-context

Post by u/FerLuisxd (title mentions u/rss.app). Config for Qwen3.8-27B on RTX 5060 Ti 16GB with vision and 85K context, 1.5GB VRAM headroom. Uses beellama (llama.cpp fork) with MTP speculative decoding. Quant IQ3_XXS-mtp from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. KV cache quant kvarn4. Author notes mmproj could be moved to CPU for more VRAM.

Qwen3.8 27B

AMD Ryzen AI 9 HX 470 96GB · llama.cpp

Tone: mixed
reported speed:
15.0 tokens/s generation · 50.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post reports stock llama.cpp on AMD Ryzen AI 9 HX 470 96GB: Qwen 3.8 27B 10-12 t/s gen, Qwen 3.8 Flash same, ~50 t/s prompt. After using strix-halo-llamacpp fork: 15-18 t/s for Flash and 12-15 t/s for 27B. generationTps set to 15 (midpoint of 12-15 for 27B after fork).

Tone: positive
reported speed:
21.5 tokens/s generation · 199.0 tokens/s prompt processing
quant:
Q5_K_M (GGUF)
kv:
f16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionlong-context

Benchmark run on Bosgame M5 (128GB/2TB) with Fedora 44, Vulkan driver 26.1.7. Used a llama.cpp fork (strix-halo-qwen4exp-b10685). MTP acceptance rate stays at 80% even at context >200K. llama-benchy results show prompt processing (pp) and generation (tg) at various context lengths: pp2048 460.38 t/s, tg512 34.86 t/s; pp4096 465.96 t/s, tg512 33.04 t/s; pp8192 437.14 t/s, tg512 31.18 t/s; pp16384 440.89 t/s, tg512 31.21 t/s; pp32768 407.48 t/s, tg512 29.65 t/s; pp65536 351.55 t/s, tg512 27.40 t/s; pp131072 265.54 t/s, tg512 23.87 t/s; pp200000 198.99 t/s, tg512 21.53 t/s. The reported promptTps and generationTps are from the pp200000 and tg512 at 200K context, respectively. The model is Qwen3.8-Flash-Next-Uncensored with Q5_K_M quant, and MTP draft model is shared-Q8_0.

Qwen3.8 Flash-Next

RTX 3090 · llama.cpp · 261,888 ctx

reported speed:
34.3 tokens/s generation · 223.7 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)
kv:
f16
flash attention:
on
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Part 4 of a series on running Qwen3.8-Flash-Next on 2x RTX 3090 with dual Broadwell Xeon and DDR4. Focus is on prefill optimization by moving expert cache off GPU during prompt processing. Reports prefill improvements of 2.2-2.5x across 8k, 37k, and 119k contexts. Decode performance unchanged. Uses llama.cpp with custom branch flashnext-e06. Hardware changed from 4 DIMMs to 6x32GB DDR4-2133 ECC. Prefill numbers: 8k 223.7 t/s (was 99.9), 37k 212.6 t/s (was 88.1), 119k 206.5 t/s (was 81.3). Decode: 8k 34.3 t/s, 37k 41.2 t/s, 119k 33.9 t/s. Time to first token improved from 82s to 37s at 8k, 424s to 176s at 37k, 1461s to 575s at 119k. Quality screen showed no regression. MTP acceptance 0.79-0.83. Uses 150-slot expert cache, 261k context, f16 KV cache. Code available at github.com/Inovello/llama.cpp/tree/flashnext-e06.

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 65,536 ctx

reported speed:
11.4 tokens/s generation · 337.9 tokens/s prompt processing
quant:
IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post title says IQ3 XXS but body also mentions 'IQ3_S - 3.4375 bpw'; model file is Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf. Generation t/s ranged 10-20 tps; detailed log shows tg=11.44 t/s with MTP speculative decoding (draft acceptance 0.4125). Prompt processing ~334-338 t/s. KV cache q4_0 for both K and V. Context 65536. More than 1GB VRAM left after loading.

Tone: positive
reported speed:
20.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User runs Qwen3.8 27B split between RTX 3060 12GB and RX 9070 XT at 20 t/s, or on 780M iGPU with 5400MHz DDR5 at 5 t/s. Prefers 5 t/s for system usability. The 20 t/s is for the split configuration; the 5 t/s is for the iGPU. The post mentions two GPUs, but the primary benchmark is the split setup.

Tone: mixed
reported speed:
38.4 tokens/s generation
quant:
IQ3_XXS
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports 38.4 t/s for Qwen3.8 27B on RTX 3060 with tuned llama.cpp, and 55.9 t/s for Qwen3.6 35B-A3B (MoE) with same tuning. Also mentions editing speeds ~188 t/s and context lengths 16K/12K for Ubuntu/WSL2. The post includes a link to a GitHub repo. The user expresses frustration about not reaching 50-60 t/s for the 27B model.

Tone: positive
reported speed:
28.0 tokens/s generation · 315.0 tokens/s prompt processing
quant:
IQ2_M
kv:
Q8
mtp (multi-token prediction):
off

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

general-conversation

Model is Qwen3.8 Flash Next, quantized IQ2_M, fits in 52GB VRAM across 4 GPUs. Generation 27-29 t/s, prefill 290-340 t/s. Context offloaded at Q8. N-gram cache on SSD. User reports ~98% clean Finnish output, but safety guardrails cause freezes in grey areas.

Tone: mixed
reported speed:
38.8 tokens/s generation · 200.0 tokens/s prompt processing
quant:
Q4_K_M (AP)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticreasoning

User reports Qwen 3.8 Flash/Next 125B MoE running on RTX 5090 with 64GB system RAM, AP quantized to Q4_K_M, using ik_llama.cpp. Decode ~38.8 tok/s, prefill ~200 tok/s, VRAM usage ~29.7 GiB. User notes it's slower and dependent on system RAM bandwidth/CPU offload. Also mentions Qwen 3.8 27B is extremely fast but weaker. Seeking advice on models and optimizations.

Qwen3.8 125B (6B active) Flash-Next

AMD MI50 32GB · llama.cpp · 130,000 ctx

reported speed:
15.0 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 15 t/s generation and 100-200 t/s prompt processing at context 130k. Uses ROCm, mentions Vulkan similar speed. Settings include flash-attn, split-mode layer, fit on.

reported speed:
25.6 tokens/s generation
quant:
Q4_K_S
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Average TG across runs at 131072 context; also tested at 196608 context with average TG 25.80 tok/s.

Qwen3.8 27B

Unknown GPU · MLX

Tone: mixed
reported speed:
29.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark comparing Qwen 3.8 27B vs Qwen 3.6 27B on oMLX. Quality improved from 81.1 to 87.7 (+8%), but speed dropped from 35 to 29 tok/s and runtime increased 5x. The post notes noticeably better quality but at the cost of tokens and time.

Qwen3.8 27B

M5 16GB · llama.cpp · 8,192 ctx

reported speed:
9.0 tokens/s generation
quant:
Q3_xxs
kv:
8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

summarization

User is new to local LLMs and asks if ~9 t/s is normal for this setup. They also ask for model recommendations for PDF summaries on 16GB RAM.

Tone: positive
reported speed:
76.7 tokens/s generation
quant:
NVFP4
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Custom Windows fork of NInfer for TP2 on two RTX 5060 Ti 16GB without P2P. Reports multiple MTP speeds: MTP0 35.75, MTP1 57.71, MTP2 63.15, MTP3 66.79, MTP4 68.5-68.9 (512-token), 70.6 (1024 tokens), 76.65 (2048 tokens).

reported speed:
15.8 tokens/s generation · 81.8 tokens/s prompt processing
quant:
4bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Comparison of MLX and llama.cpp on Qwen3.8-27B. MLX generation 15.81 tok/s, prompt 81.76 tok/s. llama.cpp generation 9.69 tok/s, prompt 99.61 tok/s. llama.cpp uses UD-Q4_K_M quant.

Tone: positive
reported speed:
65.7 tokens/s generation · 2273.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post describes PXA engine, a fork of ik_llama.cpp with vLLM plugin, for old datacenter cards. Benchmarks on 8x V100 SXM2 NVLink system. Prefill @3k: 2273 t/s (TTFT 1.4s), decode TP4 plain: 65.7 t/s. Also mentions running Qwen3.8 Flash-Next on four P100s with 150k context. Speculative decoding with k=7 gives 159.6 t/s (TP4).

Qwen3.8 125B (6B active) Flash-Next

RTX 3090 · llama.cpp · 261,888 ctx

Tone: positive
reported speed:
33.3 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)
kv:
F16
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-context

Median decode 33.3 t/s with radix-select vs 30.2 t/s old top-k at ~119k context. Quality screen: 238/240 vs 235/240 correct. No regression detected.

Qwen3.8 27B

RTX 5070 Ti · llama.cpp · 196,608 ctx

Tone: positive
reported speed:
50.0 tokens/s generation
quant:
UD-IQ4_XS (GGUF)
kv:
Q8_0 K, Q4_0 V

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports running Qwen3.8 27B at Q4 on 16GB VRAM with 200K context at 50 t/s. They pruned non-ASCII characters from embedding table and LM head to save 700MB, offloaded embedding table to save 270MB, disabled MTP, and used adaptive-kv streaming to fit 196,608 tokens. They used Q8_0 K and Q4_0 V cache quantization. The post is enthusiastic about the setup.

Qwen3.8 27B Escha-W2

RTX 4080 Super · SGLang · 98,304 ctx

Tone: positive
reported speed:
59.0 tokens/s generation
quant:
2.469 bpw
kv:
FP8 E4M3

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User pushed Qwen3.8-27B-Escha-W2 to 98K context on a 16GB 4080 Super. Reports ~59 tok/s short context, ~50.5 tok/s at 60K context. Quality surprisingly good despite aggressive 2.469 bpw quant. MTP4 gave ~67.7 tok/s at 64K context but chose no speculation for max context. Setup uses Escha's SGLang build with FP8 KV cache and BF16 SSM state.

Qwen3.8 27B

RTX 5090 · NInfer · 555,000 ctx

Tone: positive
reported speed:
117.0 tokens/s generation · 1600.0 tokens/s prompt processing
kv:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagenticcoding

Fork of NInfer with custom NVFP4 KV cache, YaRN context extension to 555k, multi-level prefix reuse with host KV safety net, tool calling improvements, and monitoring. Benchmarks: decode 117 tok/s at 400k+ ctx, cold prefill 260s for 414k tokens at 1600 tok/s, H2D restore 0.4s, host KV 30GB. Quality: LongBench matches int8, AIME 96.7%, needle-in-haystack 100%.

Tone: positive
reported speed:
10.0 tokens/s generation
quant:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

The post describes a custom AI framework running on multiple devices. The primary model mentioned is Qwen 3.8 Uncensored Q8 with 256k context on the AI PC (Strix Halo), achieving ~10 t/s. Also mentions Qwen3.6 35B MOE on RTX 5090 at ~200 t/s, and a Gemma 4 model on MacBook Air for security. The post is enthusiastic about the setup.

Qwen3.8 125B (6B active) Flash-Next

RTX 3080 20GB · ExLlamaV3 · 160,000 ctx

Tone: positive
reported speed:
25.0 tokens/s generation · 870.0 tokens/s prompt processing
quant:
4.05 EXL3

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

CPU-offloaded inference with 128GB system RAM. Compared to llama.cpp: 3.2x faster prefill, 2x faster decode. Also tested GLM 5.3 Flash with 3.05 EXL3, which ran 2x slower in decode than llama.cpp. Decode speeds warm up over time.

Qwen3.8 125B (6B active) Flash-Next

RTX 3090 · llama.cpp · 261,888 ctx

Tone: positive
reported speed:
33.3 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)
kv:
F16
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-context

Median decode improved from 30.2 to 33.3 t/s with radix-selection top-k fallback. Quality screen: 238/240 correct vs 235/240 control, no regression detected.

Tone: mixed
reported speed:
55.0 tokens/s generation
quant:
oQ4e

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingsummarization

Considering Mac Studio M5 Max 128GB for local Qwen3.8-flash-next. Mentions oQ4e quant at ~55 tok/s, also oQ5e. Debating whether to buy now or wait for M7 Ultra. Mentions use for private documents, notes, coding, general assistant work.

Qwen3.8 27B

RX 6800 16GB · llama.cpp · 131,072 ctx

reported speed:
45.0 tokens/s generation
quant:
IQ4_XS

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User currently runs llama.cpp on two RTX 2060 12GB cards (24GB total) with Qwen3.8 27B IQ4_XS at 131k context, getting ~45 tok/s. Considering upgrade to RX 6800 16GB + RX 6800 XT 16GB (32GB total) and asks about performance and ROCm/Vulkan support.

Tone: positive
reported speed:
40.0 tokens/s generation
quant:
mixed-4-8bit
kv:
8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextcoding

Generation speed ~40 tok/s on prose and 75 tok/s on coding at ~760k context. Uses 8-bit dense layers and 4-bit expert layers. Peak memory ~117GB, requires iogpu.wired_limit_mb=120000. Model weights at ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit.

reported speed:
20.3 tokens/s generation
quant:
IQ4_XS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

Benchmark comparing llama.cpp, SGLang, and FreeToken on Qwen3.8-Flash-Next. Full context TTFT: SGLang 35.4s, FreeToken 80.4s, llama.cpp+MTP 210.2s, llama.cpp baseline 258.4s. Decode at full context: SGLang 126.9, FreeToken 87.5, llama.cpp+MTP 52.6, llama.cpp baseline 20.3 tok/s. MTP improved decode 1.63x at 8K and 1.69x at 32K. Accuracy: GSM8K 95.22-95.75%, MATH-500 92.20-93.00%. Startup: llama.cpp 16s, SGLang 108s, FreeToken 126s.

Qwen3.8

RX 7900 XTX · llama.cpp

Tone: positive
reported speed:
24.0 tokens/s generation · 920.0 tokens/s prompt processing
quant:
Q3_K_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

Custom llama.cpp build optimized for 7900xtx, tensor parallel on two cards. Results for Qwen3.8 Next Q3_K_XL: 920 tk/s pp8192, 24/27 tk/s prose, 40+ tk/s code with MTP. Also tested Qwen3.8 27B Q8_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code. Qwen3.6 27B Q4_K_M single card: 1020 tk/s pp8192, 58/60 tk/s prose, 75/80 tk/s code.

Qwen3.8 27B

Unknown GPU

Tone: positive

Post describes a task-aware quantization (TAK) of Qwen3.8-27B achieving 82.81% on a reasoning benchmark vs 77.34% for Unsloth UD IQ2_S and 83.59% for BF16. No hardware or engine mentioned. The model is a quantized variant, but no quant code is given. The post is enthusiastic about the method's results.

Tone: mixed
reported speed:
150.0 tokens/s generation · 7000.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User switched from Qwen 3.6 27B to Qwen 3.8 Next Flash. Reports very verbose output with long thinking times (13 minutes on single-turn coding requests). Mentions using pi.dev and BYOK to VSCode. Criticizes output quality for decision-making tasks, calling it 'alphabet soup'.

Tone: positive
reported speed:
20.0 tokens/s generation · 800.0 tokens/s prompt processing
quant:
IQ3_K_XXS
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

Model is Qwen 3.8 27B IQ3_K_XXS by Unsloth. Fits fully on 4060Ti 16GB with ~100k context at Q8 KV cache, dropping mmproj and MTP. Average 800 tk/s prefill and 20 tk/s decode (17 tk/s after 64k context). Used for agentic coding with parallel tool calls; successfully merged a feature branch. Author is impressed and plans to upgrade to R9700.

Qwen3.8 27B

M5 Ultra 96GB

reported speed:
15.0 tokens/s generation
quant:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User is comparing M5 Ultra 96GB vs M5 Max 128GB for running Qwen3.8-27B at Q8, currently getting ~15 tok/s on the Ultra. They also discuss Qwen3.8-Flash-Next, a multimodal MoE with 176B total params and ~6B active, which they estimate won't fit in 96GB. They mention MLX and llama.cpp as potential engines but don't confirm using them.

Qwen3.8 27B

Unknown GPU

vision

Benchmark of vision models for calorie estimation from meal photos. Qwen 3.8 27b scored 16% within 20% error, mean bias +64 kcal, median error 148 kcal. Other models tested include GLM 5.3 Flash, Qwen 3.8 Max, Muse Glimmer 30b, Qwen 3.8 Flash, DeepSeek v4 Flash Vision, and Muse Spark 1.3. The user notes that model size does not correlate with performance and that the best model on consumer hardware (~32GB VRAM) depends on the task.

Qwen3.8 2400B (95B active)

RTX 5090 · llama.cpp · 65,536 ctx

Tone: negative
reported speed:
0.3 tokens/s generation
quant:
UD-Q1_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Attempted to run Qwen3.8 2.4T on dual RTX 5090s and three RTX 3090s with 96GB system RAM. Model size 397GB, far exceeds available VRAM+RAM. Generation speed 0.25 t/s, deemed unusable. Context length 64k.

Qwen3.8 2400B (95B active)

RTX 5090 · llama.cpp · 512 ctx

Tone: positive
reported speed:
0.8 tokens/s generation · 0.8 tokens/s prompt processing
quant:
Q1_0 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Ran Qwen3.8-2.4T-A95B with Unsloth GGUF Q1_0 (397 GiB) on RTX 5090 + RTX 5060 Ti using llama.cpp (Unsloth build 10360). Enabled native MTP speculative decoding with n_max=3, p_min=0.5, and block 92 experts on CPU. Achieved ~0.80 tok/s generation. MTP acceptance 90.48%, +3.64% throughput vs no MTP. VRAM usage: 5090 29.6GB, 5060 Ti 12.3GB.

Qwen3.8 max

Unknown GPU

Tone: mixed
reported speed:
65.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User tested Qwen 3.8 max (preview) and reports generation speed of 60-65 t/s. Praises creativity and writing but criticizes excessive thinking time and hallucination.

Tone: positive
reported speed:
8.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks about MTP or DFlash head for Qwen 3.8 27B. Mentions 35B-A3B as daily driver. Reports ~8 tok/s for Qwen 3.6 27B with MTP on 32GB unified memory.

Tone: positive
reported speed:
200.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Day-0 support for Qwen3.8-27B in NInfer. Reports ~200 tok/s generation with speculative decoding on a single RTX 5090. Mentions improvements: up to 8 concurrent requests, shared paged KV cache, ReplaySSM for GDN, PDL usage.

Qwen3.8 27B

RTX Pro 6000 Blackwell · vLLM · 262,144 ctx

Tone: mixed
reported speed:
114.8 tokens/s generation
quant:
FP8
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Comparison of Qwen3.8-27B-FP8 vs Qwen3.6-27B-FP8 on RTX PRO 6000 Blackwell with vLLM. MTP sweep results: Qwen3.8 peaked at MTP 6 with 114.8 tok/s, while Qwen3.6 peaked at MTP 7 with 131.6 tok/s. Qwen3.8 is 5-20% slower across MTP steps but quality is comparable or better at some steps. Pre-sweep manual run at MTP 5 showed Qwen3.8 faster (108.9 vs 103.6 tok/s).

Qwen3.8 27B

RTX 5070 Ti Laptop 12GB · Unsloth Studio · 8,192 ctx

Tone: positive
reported speed:
4.5 tokens/s generation · 23.5 tokens/s prompt processing
quant:
UD-Q4_K_XL
kv:
Q8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

MTP acceptance 78-83%. Longer response 3.26 t/s, short factual 4.42 t/s, coding 4.53 t/s. Prompt processing 20-27 t/s. Model size ~17.9GB, offloading to CPU.

Tone: positive
reported speed:
200.0 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Day-0 support for Qwen3.8-27B with NVFP4 + DSpark. 200+ tok/s decode on RTX 5090 and RTX Pro 6000; 38 tok/s on DGX Spark. Also mentions H200 in the cookbook link.

Qwen3.8 27B

RTX 3090 · llama.cpp

Tone: positive
reported speed:
35.0 tokens/s generation
quant:
IQ4_NL
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 34-36 t/s on RTX 3090 with Qwen3.8 27B using llama.cpp, IQ4_NL quant, KV cache at Q8. Compares to 50-60 t/s on Qwen3.6 27B. Praises model as 'next gen'.

Qwen3.8 27B

RTX 5090 · llama.cpp · 32,768 ctx

Tone: mixed
reported speed:
58.1 tokens/s generation · 479.2 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

creative-writing

Compared Qwen3.8-27B vs Qwen3.6-27B on RTX 5090. Subjective quality improvement smaller than benchmarks suggest. Medium reasoning used fewer tokens than low on 3/5 prompts. Concurrency knee at 3 requests for Qwen3.8 vs ~7 for Qwen3.6. Prefill and decode speeds at concurrency 1 reported.

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
40.0 tokens/s generation
quant:
Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

Reported ~40 tok/s with Q4_K_XL quant on 2x RTX 3060 12GB using llama.cpp with flash attention and speculative decoding. Model successfully wrote CUDA code but couldn't run due to VRAM limits.

Tone: mixed
reported speed:
72.1 tokens/s generation
quant:
Q4_K_M

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

summarization

Compared Qwen3.8-27B vs Qwen3.6-27B on fact-extraction task. Qwen3.8 scored 0.7030 F1 vs Qwen3.6's 0.7177, a statistical tie. Decode throughput fell from 85.6 to 72.1 t/s (~16%) under closest saved configs, though different llama.cpp builds. Qwen3.8 produced shorter answers, lowering end-to-end latency. Author expected larger gains based on public benchmarks; sees small gains on task-specific tests but massive gains only on benchmarks the model was trained on.

Qwen3.8 27B

RTX 3090 · llama.cpp · 200,000 ctx

Tone: mixed
reported speed:
58.5 tokens/s generation
quant:
Q8_K_XL (GGUF)
kv:
f16
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

Dual RTX 3090 (24GB each, 48GB total). Qwen3.8-27B with Q8_K_XL quant (31.5GB) and mmproj-F16 (0.93GB). Context 200K with f16 KV cache. Generation speeds: 73 tok/s on code, 58.5 tok/s on prose (median of 3 draws, temp 0). MTP draft depth 2 gives best acceptance (92% code, 68% prose). Hybrid architecture: 48 of 64 layers are Gated DeltaNet (linear attention), only 16 full attention layers plus MTP head = 17 KV-caching layers. Cold start 32s. Tensor split 50/50 but ~1.1GB lopsided causing OOM risk at 220K context. User asks for optimization suggestions for dual 3090 setups.

Tone: positive
reported speed:
70.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

NInfer is a custom C++/CUDA inference runtime. The 3090 port targets Ampere GPUs. Results are sustained runs with 1024 output tokens and CUDA Graphs enabled. Also mentions Qwen3-35B-A3B running at ~260 tok/s single-stream and over 400 tok/s on repetitive workloads.