- reported speed:
- 70.0 tokens/s generation · 1300.0 tokens/s prompt processing
- quant:
- W4A16 (AutoRound)
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingcreative-writing
User reports Qwen3.8 Flash-Next at 70 t/s generation and 1.3k t/s prefill on 4x V620 (128GB total VRAM) at 128k+ context.
Setup is a vLLM fork with AutoRound W4A16 quantization and MTP-2 speculative decoding, on an EPYC 7452 with 256GB DDR4. Prose generation runs at 60 t/s.
Power draw is 700-900W during prefill and 500-600W during decode. User was initially disappointed with Qwen3.8 27B speeds but is satisfied with Flash-Next.
- reported speed:
- 70.0 tokens/s generation · 1386.0 tokens/s prompt processing
- quant:
- W4A16 (AutoRound)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codinglong-context
User reports Qwen3.8 Flash-Next at 70.0 t/s generation and 1,386 t/s prompt processing on 4 Radeon Pro V620 GPUs at 32K context on the coding suite.
Setup is vLLM with the Intel/Qwen3.8-Flash-Next-W4A16-AutoRound quant, using a custom RDNA vLLM fork.
At 64K context the coding suite reached 72.5 t/s generation and 1,369 t/s prompt processing; at 128K it reached 68.3 t/s and 1,296 t/s. The regular suite measured 57.9, 59.5, and 56.1 t/s generation at 32K, 64K, and 128K respectively.
- reported speed:
- 35.4 tokens/s generation · 355.4 tokens/s prompt processing
- quant:
- Q6_K_XL (GGUF)
- kv:
- f16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Muse Glimmer 30B on AMD v620 GPUs, with Q6 on one GPU at 355.38 t/s prefill and 35.38 t/s generation.
Setup is a tensor split across two GPUs: Q6 gives 472.27 t/s prefill and 36.32 t/s generation, Q8 gives 550.04 t/s prefill and 26.55 t/s generation.
Speculative decoding uses a dflash draft model.
- reported speed:
- 21.1 tokens/s generation · 276.2 tokens/s prompt processing
- quant:
- IQ3_XXS
- kv:
- FP16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek-V4-Flash-0731 at 21.06 t/s continuous generation and 30.59 t/s sustained accepted generation on 4x AMD Radeon Pro V620 with 128 GB total.
Setup is the unsloth IQ3_XXS quant with speculative decoding using a DSpark Q8 drafter on CPU RAM and 2 draft tokens. Prompt ingestion is 276.23 t/s at 32K and 379.01 t/s at 4K.
The user calls the model at q3_xxs "fucking stupid" and lobotomized, and notes there is not enough VRAM for the drafter, which affects TG speeds.
- reported speed:
- 30.0 tokens/s generation · 1000.0 tokens/s prompt processing
- quant:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at 25.6 t/s on a Threadripper 3975WX with 4x V620 GPUs.
Setup is llama.cpp with the Huihui abliterated Q4_K_S GGUF and a Q4 KV cache, the whole 24 GB card in use.
A second run at 196,608 context averaged 25.8 t/s. The user also mentions Qwen3.6-35B with 3.5k prefill at 70-80 t/s, but the primary model is Qwen3.8-27B-q8.
- reported speed:
- 14.3 tokens/s generation · 364.0 tokens/s prompt processing
- quant:
- Q5_K_P (gguf)
- kv:
- Q4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Qwen3.8-27B and Gemma4-26B-A4B on an AMD V620 using llama.cpp with ROCm and Vulkan.
Qwen results are shown; Gemma results appear in a table but are not extracted as the primary model.