- reported speed:
- 15.8 tokens/s generation · 81.8 tokens/s prompt processing
- quant:
- 4bit (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Benchmark compares MLX vs llama.cpp on Qwen3.8-27B. MLX: mlx-community/Qwen3.8-27B-4bit, ~16.1GB, prompt 81.76 tok/s, generation 15.81 tok/s, peak memory 16.39GB. llama.cpp: unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M, 15.32 GiB / 27.32B params, full Metal offload, Flash Attention enabled, prompt 99.61 ± 0.44 tok/s, generation 9.69 ± 0.34 tok/s. llama.cpp ~22% faster prompt processing, MLX ~63% faster generation.
- reported speed:
- 18.3 tokens/s generation · 235.7 tokens/s prompt processing
- quant:
- oQ4e (MLX)
- kv:
- 8.0
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports slow performance with Qwen3.8-Flash-Next on M2 Max 96GB using oMLX engine. Benchmark results show pp1024/tg128: 124.1 pp TPS, 20.4 tg TPS; pp4096/tg128: 170.7 pp TPS, 16.7 tg TPS; pp8192/tg128: 214.3 pp TPS, 18.5 tg TPS; pp16384/tg128: 235.7 pp TPS, 18.3 tg TPS. Also tried oQ4e-fp16-mtp, oQ3-fp16-mtp, oQ3-MTP variants. fp16 variants even slower (~150pp, 8tg). User notes llama.cpp achieves 350-400 pp TPS. Model is MoE with 6B active parameters. Context length set to 32768. KV cache quant is turboquant_kv_bits 8.0. Engine is oMLX (custom MLX-based).
- reported speed:
- 38.0 tokens/s generation
- quant:
- T5 (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingmath
Runs on a custom oMLX fork with a Metal kernel for ternary experts; stock oMLX/mlx-lm won't load it. ~35.6 GiB resident: routed expert gate/up as ternary (Bonsai T5 packing, ~1.875 bpw), expert down projections as Q3, 53 GB n-gram table left on SSD as Q8 and mmapped per token. All low-bit tensors fitted with Unsloth's imatrix. Prefill 210-380 tok/s. 64K context confirmed (65,536-token prompt + 256 output at 30.7 tok/s, 42.3 GiB physical peak); 96K trips the prefill guard. Physical peak at 8K context ~41.5 GiB; swaps ~2 GiB once at load. Setup: oMLX memory guard 'safe' profile, limit 48 GB, one model, one request at a time. Quality vs Unsloth UD-Q4_K_XL (doesn't fit in 48 GB): KLD vs Q8_0 0.49 vs 0.036; MMLU 83.0% vs 89.7%; GSM8K 90.0% vs 92.0%; HumanEval 92.7% vs 95.7%. Known wart: sometimes ignores 'answer with just the letter' in Chinese.
- reported speed:
- 45.0 tokens/s generation · 300.0 tokens/s prompt processing
- quant:
- IQ3_XXS (GGUF)
- kv:
- kvarn4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
visionlong-context
Post by u/FerLuisxd (title mentions u/rss.app). Config for Qwen3.8-27B on RTX 5060 Ti 16GB with vision and 85K context, 1.5GB VRAM headroom. Uses beellama (llama.cpp fork) with MTP speculative decoding. Quant IQ3_XXS-mtp from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. KV cache quant kvarn4. Author notes mmproj could be moved to CPU for more VRAM.
- reported speed:
- 15.0 tokens/s generation · 50.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Post reports stock llama.cpp on AMD Ryzen AI 9 HX 470 96GB: Qwen 3.8 27B 10-12 t/s gen, Qwen 3.8 Flash same, ~50 t/s prompt. After using strix-halo-llamacpp fork: 15-18 t/s for Flash and 12-15 t/s for 27B. generationTps set to 15 (midpoint of 12-15 for 27B after fork).
- reported speed:
- 21.5 tokens/s generation · 199.0 tokens/s prompt processing
- quant:
- Q5_K_M (GGUF)
- kv:
- f16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
visionlong-context
Benchmark run on Bosgame M5 (128GB/2TB) with Fedora 44, Vulkan driver 26.1.7. Used a llama.cpp fork (strix-halo-qwen4exp-b10685). MTP acceptance rate stays at 80% even at context >200K. llama-benchy results show prompt processing (pp) and generation (tg) at various context lengths: pp2048 460.38 t/s, tg512 34.86 t/s; pp4096 465.96 t/s, tg512 33.04 t/s; pp8192 437.14 t/s, tg512 31.18 t/s; pp16384 440.89 t/s, tg512 31.21 t/s; pp32768 407.48 t/s, tg512 29.65 t/s; pp65536 351.55 t/s, tg512 27.40 t/s; pp131072 265.54 t/s, tg512 23.87 t/s; pp200000 198.99 t/s, tg512 21.53 t/s. The reported promptTps and generationTps are from the pp200000 and tg512 at 200K context, respectively. The model is Qwen3.8-Flash-Next-Uncensored with Q5_K_M quant, and MTP draft model is shared-Q8_0.
- reported speed:
- 34.3 tokens/s generation · 223.7 tokens/s prompt processing
- quant:
- UD-Q4_K_XL (GGUF)
- kv:
- f16
- flash attention:
- on
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Part 4 of a series on running Qwen3.8-Flash-Next on 2x RTX 3090 with dual Broadwell Xeon and DDR4. Focus is on prefill optimization by moving expert cache off GPU during prompt processing. Reports prefill improvements of 2.2-2.5x across 8k, 37k, and 119k contexts. Decode performance unchanged. Uses llama.cpp with custom branch flashnext-e06. Hardware changed from 4 DIMMs to 6x32GB DDR4-2133 ECC. Prefill numbers: 8k 223.7 t/s (was 99.9), 37k 212.6 t/s (was 88.1), 119k 206.5 t/s (was 81.3). Decode: 8k 34.3 t/s, 37k 41.2 t/s, 119k 33.9 t/s. Time to first token improved from 82s to 37s at 8k, 424s to 176s at 37k, 1461s to 575s at 119k. Quality screen showed no regression. MTP acceptance 0.79-0.83. Uses 150-slot expert cache, 261k context, f16 KV cache. Code available at github.com/Inovello/llama.cpp/tree/flashnext-e06.
- reported speed:
- 11.4 tokens/s generation · 337.9 tokens/s prompt processing
- quant:
- IQ3_XXS (GGUF)
- kv:
- q4_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Post title says IQ3 XXS but body also mentions 'IQ3_S - 3.4375 bpw'; model file is Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf. Generation t/s ranged 10-20 tps; detailed log shows tg=11.44 t/s with MTP speculative decoding (draft acceptance 0.4125). Prompt processing ~334-338 t/s. KV cache q4_0 for both K and V. Context 65536. More than 1GB VRAM left after loading.
- reported speed:
- 20.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User runs Qwen3.8 27B split between RTX 3060 12GB and RX 9070 XT at 20 t/s, or on 780M iGPU with 5400MHz DDR5 at 5 t/s. Prefers 5 t/s for system usability. The 20 t/s is for the split configuration; the 5 t/s is for the iGPU. The post mentions two GPUs, but the primary benchmark is the split setup.
- reported speed:
- 38.4 tokens/s generation
- quant:
- IQ3_XXS
- kv:
- Q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
User reports 38.4 t/s for Qwen3.8 27B on RTX 3060 with tuned llama.cpp, and 55.9 t/s for Qwen3.6 35B-A3B (MoE) with same tuning. Also mentions editing speeds ~188 t/s and context lengths 16K/12K for Ubuntu/WSL2. The post includes a link to a GitHub repo. The user expresses frustration about not reaching 50-60 t/s for the 27B model.
- reported speed:
- 28.0 tokens/s generation · 315.0 tokens/s prompt processing
- quant:
- IQ2_M
- kv:
- Q8
- mtp (multi-token prediction):
- off
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
general-conversation
Model is Qwen3.8 Flash Next, quantized IQ2_M, fits in 52GB VRAM across 4 GPUs. Generation 27-29 t/s, prefill 290-340 t/s. Context offloaded at Q8. N-gram cache on SSD. User reports ~98% clean Finnish output, but safety guardrails cause freezes in grey areas.
- reported speed:
- 38.8 tokens/s generation · 200.0 tokens/s prompt processing
- quant:
- Q4_K_M (AP)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingagenticreasoning
User reports Qwen 3.8 Flash/Next 125B MoE running on RTX 5090 with 64GB system RAM, AP quantized to Q4_K_M, using ik_llama.cpp. Decode ~38.8 tok/s, prefill ~200 tok/s, VRAM usage ~29.7 GiB. User notes it's slower and dependent on system RAM bandwidth/CPU offload. Also mentions Qwen 3.8 27B is extremely fast but weaker. Seeking advice on models and optimizations.
- reported speed:
- 15.0 tokens/s generation
- quant:
- UD-Q4_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 15 t/s generation and 100-200 t/s prompt processing at context 130k. Uses ROCm, mentions Vulkan similar speed. Settings include flash-attn, split-mode layer, fit on.
- reported speed:
- 25.6 tokens/s generation
- quant:
- Q4_K_S
- kv:
- Q4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Average TG across runs at 131072 context; also tested at 196608 context with average TG 25.80 tok/s.
- reported speed:
- 29.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Benchmark comparing Qwen 3.8 27B vs Qwen 3.6 27B on oMLX. Quality improved from 81.1 to 87.7 (+8%), but speed dropped from 35 to 29 tok/s and runtime increased 5x. The post notes noticeably better quality but at the cost of tokens and time.
- reported speed:
- 9.0 tokens/s generation
- quant:
- Q3_xxs
- kv:
- 8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
summarization
User is new to local LLMs and asks if ~9 t/s is normal for this setup. They also ask for model recommendations for PDF summaries on 16GB RAM.
- reported speed:
- 76.7 tokens/s generation
- quant:
- NVFP4
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Custom Windows fork of NInfer for TP2 on two RTX 5060 Ti 16GB without P2P. Reports multiple MTP speeds: MTP0 35.75, MTP1 57.71, MTP2 63.15, MTP3 66.79, MTP4 68.5-68.9 (512-token), 70.6 (1024 tokens), 76.65 (2048 tokens).
- reported speed:
- 15.8 tokens/s generation · 81.8 tokens/s prompt processing
- quant:
- 4bit
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Comparison of MLX and llama.cpp on Qwen3.8-27B. MLX generation 15.81 tok/s, prompt 81.76 tok/s. llama.cpp generation 9.69 tok/s, prompt 99.61 tok/s. llama.cpp uses UD-Q4_K_M quant.
- reported speed:
- 65.7 tokens/s generation · 2273.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Post describes PXA engine, a fork of ik_llama.cpp with vLLM plugin, for old datacenter cards. Benchmarks on 8x V100 SXM2 NVLink system. Prefill @3k: 2273 t/s (TTFT 1.4s), decode TP4 plain: 65.7 t/s. Also mentions running Qwen3.8 Flash-Next on four P100s with 150k context. Speculative decoding with k=7 gives 159.6 t/s (TP4).
- reported speed:
- 33.3 tokens/s generation
- quant:
- UD-Q4_K_XL (GGUF)
- kv:
- F16
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
long-context
Median decode 33.3 t/s with radix-select vs 30.2 t/s old top-k at ~119k context. Quality screen: 238/240 vs 235/240 correct. No regression detected.
- reported speed:
- 50.0 tokens/s generation
- quant:
- UD-IQ4_XS (GGUF)
- kv:
- Q8_0 K, Q4_0 V
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
User reports running Qwen3.8 27B at Q4 on 16GB VRAM with 200K context at 50 t/s. They pruned non-ASCII characters from embedding table and LM head to save 700MB, offloaded embedding table to save 270MB, disabled MTP, and used adaptive-kv streaming to fit 196,608 tokens. They used Q8_0 K and Q4_0 V cache quantization. The post is enthusiastic about the setup.
- reported speed:
- 59.0 tokens/s generation
- quant:
- 2.469 bpw
- kv:
- FP8 E4M3
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingagenticlong-context
User pushed Qwen3.8-27B-Escha-W2 to 98K context on a 16GB 4080 Super. Reports ~59 tok/s short context, ~50.5 tok/s at 60K context. Quality surprisingly good despite aggressive 2.469 bpw quant. MTP4 gave ~67.7 tok/s at 64K context but chose no speculation for max context. Setup uses Escha's SGLang build with FP8 KV cache and BF16 SSM state.
- reported speed:
- 117.0 tokens/s generation · 1600.0 tokens/s prompt processing
- kv:
- NVFP4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
long-contextagenticcoding
Fork of NInfer with custom NVFP4 KV cache, YaRN context extension to 555k, multi-level prefix reuse with host KV safety net, tool calling improvements, and monitoring. Benchmarks: decode 117 tok/s at 400k+ ctx, cold prefill 260s for 414k tokens at 1600 tok/s, H2D restore 0.4s, host KV 30GB. Quality: LongBench matches int8, AIME 96.7%, needle-in-haystack 100%.
- reported speed:
- 10.0 tokens/s generation
- quant:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
The post describes a custom AI framework running on multiple devices. The primary model mentioned is Qwen 3.8 Uncensored Q8 with 256k context on the AI PC (Strix Halo), achieving ~10 t/s. Also mentions Qwen3.6 35B MOE on RTX 5090 at ~200 t/s, and a Gemma 4 model on MacBook Air for security. The post is enthusiastic about the setup.
- reported speed:
- 25.0 tokens/s generation · 870.0 tokens/s prompt processing
- quant:
- 4.05 EXL3
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
CPU-offloaded inference with 128GB system RAM. Compared to llama.cpp: 3.2x faster prefill, 2x faster decode. Also tested GLM 5.3 Flash with 3.05 EXL3, which ran 2x slower in decode than llama.cpp. Decode speeds warm up over time.
- reported speed:
- 33.3 tokens/s generation
- quant:
- UD-Q4_K_XL (GGUF)
- kv:
- F16
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
long-context
Median decode improved from 30.2 to 33.3 t/s with radix-selection top-k fallback. Quality screen: 238/240 correct vs 235/240 control, no regression detected.
- reported speed:
- 55.0 tokens/s generation
- quant:
- oQ4e
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingsummarization
Considering Mac Studio M5 Max 128GB for local Qwen3.8-flash-next. Mentions oQ4e quant at ~55 tok/s, also oQ5e. Debating whether to buy now or wait for M7 Ultra. Mentions use for private documents, notes, coding, general assistant work.
- reported speed:
- 45.0 tokens/s generation
- quant:
- IQ4_XS
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User currently runs llama.cpp on two RTX 2060 12GB cards (24GB total) with Qwen3.8 27B IQ4_XS at 131k context, getting ~45 tok/s. Considering upgrade to RX 6800 16GB + RX 6800 XT 16GB (32GB total) and asks about performance and ROCm/Vulkan support.
- reported speed:
- 40.0 tokens/s generation
- quant:
- mixed-4-8bit
- kv:
- 8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
long-contextcoding
Generation speed ~40 tok/s on prose and 75 tok/s on coding at ~760k context. Uses 8-bit dense layers and 4-bit expert layers. Peak memory ~117GB, requires iogpu.wired_limit_mb=120000. Model weights at ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit.
- reported speed:
- 20.3 tokens/s generation
- quant:
- IQ4_XS (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codinglong-context
Benchmark comparing llama.cpp, SGLang, and FreeToken on Qwen3.8-Flash-Next. Full context TTFT: SGLang 35.4s, FreeToken 80.4s, llama.cpp+MTP 210.2s, llama.cpp baseline 258.4s. Decode at full context: SGLang 126.9, FreeToken 87.5, llama.cpp+MTP 52.6, llama.cpp baseline 20.3 tok/s. MTP improved decode 1.63x at 8K and 1.69x at 32K. Accuracy: GSM8K 95.22-95.75%, MATH-500 92.20-93.00%. Startup: llama.cpp 16s, SGLang 108s, FreeToken 126s.
- reported speed:
- 24.0 tokens/s generation · 920.0 tokens/s prompt processing
- quant:
- Q3_K_XL
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
Custom llama.cpp build optimized for 7900xtx, tensor parallel on two cards. Results for Qwen3.8 Next Q3_K_XL: 920 tk/s pp8192, 24/27 tk/s prose, 40+ tk/s code with MTP. Also tested Qwen3.8 27B Q8_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code. Qwen3.6 27B Q4_K_M single card: 1020 tk/s pp8192, 58/60 tk/s prose, 75/80 tk/s code.
Post describes a task-aware quantization (TAK) of Qwen3.8-27B achieving 82.81% on a reasoning benchmark vs 77.34% for Unsloth UD IQ2_S and 83.59% for BF16. No hardware or engine mentioned. The model is a quantized variant, but no quant code is given. The post is enthusiastic about the method's results.
- reported speed:
- 150.0 tokens/s generation · 7000.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
User switched from Qwen 3.6 27B to Qwen 3.8 Next Flash. Reports very verbose output with long thinking times (13 minutes on single-turn coding requests). Mentions using pi.dev and BYOK to VSCode. Criticizes output quality for decision-making tasks, calling it 'alphabet soup'.
- reported speed:
- 20.0 tokens/s generation · 800.0 tokens/s prompt processing
- quant:
- IQ3_K_XXS
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingagentic
Model is Qwen 3.8 27B IQ3_K_XXS by Unsloth. Fits fully on 4060Ti 16GB with ~100k context at Q8 KV cache, dropping mmproj and MTP. Average 800 tk/s prefill and 20 tk/s decode (17 tk/s after 64k context). Used for agentic coding with parallel tool calls; successfully merged a feature branch. Author is impressed and plans to upgrade to R9700.
- reported speed:
- 15.0 tokens/s generation
- quant:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
agentic
User is comparing M5 Ultra 96GB vs M5 Max 128GB for running Qwen3.8-27B at Q8, currently getting ~15 tok/s on the Ultra. They also discuss Qwen3.8-Flash-Next, a multimodal MoE with 176B total params and ~6B active, which they estimate won't fit in 96GB. They mention MLX and llama.cpp as potential engines but don't confirm using them.
vision
Benchmark of vision models for calorie estimation from meal photos. Qwen 3.8 27b scored 16% within 20% error, mean bias +64 kcal, median error 148 kcal. Other models tested include GLM 5.3 Flash, Qwen 3.8 Max, Muse Glimmer 30b, Qwen 3.8 Flash, DeepSeek v4 Flash Vision, and Muse Spark 1.3. The user notes that model size does not correlate with performance and that the best model on consumer hardware (~32GB VRAM) depends on the task.
- reported speed:
- 0.3 tokens/s generation
- quant:
- UD-Q1_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Attempted to run Qwen3.8 2.4T on dual RTX 5090s and three RTX 3090s with 96GB system RAM. Model size 397GB, far exceeds available VRAM+RAM. Generation speed 0.25 t/s, deemed unusable. Context length 64k.
- reported speed:
- 0.8 tokens/s generation · 0.8 tokens/s prompt processing
- quant:
- Q1_0 (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Ran Qwen3.8-2.4T-A95B with Unsloth GGUF Q1_0 (397 GiB) on RTX 5090 + RTX 5060 Ti using llama.cpp (Unsloth build 10360). Enabled native MTP speculative decoding with n_max=3, p_min=0.5, and block 92 experts on CPU. Achieved ~0.80 tok/s generation. MTP acceptance 90.48%, +3.64% throughput vs no MTP. VRAM usage: 5090 29.6GB, 5060 Ti 12.3GB.
- reported speed:
- 65.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User tested Qwen 3.8 max (preview) and reports generation speed of 60-65 t/s. Praises creativity and writing but criticizes excessive thinking time and hallucination.
- reported speed:
- 8.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User asks about MTP or DFlash head for Qwen 3.8 27B. Mentions 35B-A3B as daily driver. Reports ~8 tok/s for Qwen 3.6 27B with MTP on 32GB unified memory.
- reported speed:
- 200.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Day-0 support for Qwen3.8-27B in NInfer. Reports ~200 tok/s generation with speculative decoding on a single RTX 5090. Mentions improvements: up to 8 concurrent requests, shared paged KV cache, ReplaySSM for GDN, PDL usage.
- reported speed:
- 114.8 tokens/s generation
- quant:
- FP8
- kv:
- fp8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Comparison of Qwen3.8-27B-FP8 vs Qwen3.6-27B-FP8 on RTX PRO 6000 Blackwell with vLLM. MTP sweep results: Qwen3.8 peaked at MTP 6 with 114.8 tok/s, while Qwen3.6 peaked at MTP 7 with 131.6 tok/s. Qwen3.8 is 5-20% slower across MTP steps but quality is comparable or better at some steps. Pre-sweep manual run at MTP 5 showed Qwen3.8 faster (108.9 vs 103.6 tok/s).
- reported speed:
- 4.5 tokens/s generation · 23.5 tokens/s prompt processing
- quant:
- UD-Q4_K_XL
- kv:
- Q8
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
MTP acceptance 78-83%. Longer response 3.26 t/s, short factual 4.42 t/s, coding 4.53 t/s. Prompt processing 20-27 t/s. Model size ~17.9GB, offloading to CPU.
- reported speed:
- 200.0 tokens/s generation
- quant:
- NVFP4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Day-0 support for Qwen3.8-27B with NVFP4 + DSpark. 200+ tok/s decode on RTX 5090 and RTX Pro 6000; 38 tok/s on DGX Spark. Also mentions H200 in the cookbook link.
- reported speed:
- 35.0 tokens/s generation
- quant:
- IQ4_NL
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 34-36 t/s on RTX 3090 with Qwen3.8 27B using llama.cpp, IQ4_NL quant, KV cache at Q8. Compares to 50-60 t/s on Qwen3.6 27B. Praises model as 'next gen'.
- reported speed:
- 58.1 tokens/s generation · 479.2 tokens/s prompt processing
- quant:
- Q6_K (GGUF)
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
creative-writing
Compared Qwen3.8-27B vs Qwen3.6-27B on RTX 5090. Subjective quality improvement smaller than benchmarks suggest. Medium reasoning used fewer tokens than low on 3/5 prompts. Concurrency knee at 3 requests for Qwen3.8 vs ~7 for Qwen3.6. Prefill and decode speeds at concurrency 1 reported.
- reported speed:
- 40.0 tokens/s generation
- quant:
- Q4_K_XL (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
Reported ~40 tok/s with Q4_K_XL quant on 2x RTX 3060 12GB using llama.cpp with flash attention and speculative decoding. Model successfully wrote CUDA code but couldn't run due to VRAM limits.
- reported speed:
- 72.1 tokens/s generation
- quant:
- Q4_K_M
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
summarization
Compared Qwen3.8-27B vs Qwen3.6-27B on fact-extraction task. Qwen3.8 scored 0.7030 F1 vs Qwen3.6's 0.7177, a statistical tie. Decode throughput fell from 85.6 to 72.1 t/s (~16%) under closest saved configs, though different llama.cpp builds. Qwen3.8 produced shorter answers, lowering end-to-end latency. Author expected larger gains based on public benchmarks; sees small gains on task-specific tests but massive gains only on benchmarks the model was trained on.
- reported speed:
- 58.5 tokens/s generation
- quant:
- Q8_K_XL (GGUF)
- kv:
- f16
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
Dual RTX 3090 (24GB each, 48GB total). Qwen3.8-27B with Q8_K_XL quant (31.5GB) and mmproj-F16 (0.93GB). Context 200K with f16 KV cache. Generation speeds: 73 tok/s on code, 58.5 tok/s on prose (median of 3 draws, temp 0). MTP draft depth 2 gives best acceptance (92% code, 68% prose). Hybrid architecture: 48 of 64 layers are Gated DeltaNet (linear attention), only 16 full attention layers plus MTP head = 17 KV-caching layers. Cold start 32s. Tensor split 50/50 but ~1.1GB lopsided causing OOM risk at 220K context. User asks for optimization suggestions for dual 3090 setups.
- reported speed:
- 70.2 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
NInfer is a custom C++/CUDA inference runtime. The 3090 port targets Ampere GPUs. Results are sustained runs with 1024 output tokens and CUDA Graphs enabled. Also mentions Qwen3-35B-A3B running at ~260 tok/s single-stream and over 400 tok/s on repetitive workloads.