llamaperf

Qwen3-Next Flash

on 6× BC-250 · llama.cpp · 100,000 ctx

Tone: positive
Oct 2, 2026
Throughput
28.0 t/s gen
Quant
IQ2_XS (GGUF)

Summary

User reports Qwen3-Next Flash IQ2_XS at around 28 tok/s for short generation on a cluster of 4 BC-250 boards, dropping to 24 tok/s at 50k context with around 115 t/s prefill. Setup is llama.cpp with Vulkan and RPC over 1Gb Ethernet, 100k context, across 6 BC-250 ex-mining boards in an ASRock 4U12G case. The other two boards run Qwen3.6 35B Q4 at 60 tok/s with 100k context and 450 t/s prefill.