llamaperf

Qwen3.8 125B (6B active) Flash-Next

on 4× NVIDIA V100 16GB · Strata · 252,000 ctx

Tone: positive
Oct 4, 2026
Throughput
7090.0 t/s pp
Quant
UD-Q4_K_XL (GGUF)
VRAM reported
16 GB

Summary

User reports Qwen3.8-Flash-Next UD-Q4_K_XL at up to 7,350 tok/s prefill and ~113 tok/s decode on an IBM AC922 with 4x Tesla V100 16GB SXM2 GPUs. Setup is a forked Strata engine with MTP speculative decoding, 72 GiB of experts page-locked in RAM across both sockets, and GPUs pulling from NVLink 2.0 at ~70 GB/s each. The 252K-token prompt prefilled at 7,090 tok/s in 35 s; follow-up at that depth gave first token after 0.26 s and 60 tok/s. Generation varied by workload: ~113 tok/s peak on JSON, ~100 on code, ~84 on prose. Stock llama.cpp on the same machine produced 130 tok/s prefill and 15 tok/s decode.