Qwen3-Next Flash
on 6× BC-250 · llama.cpp · 100,000 ctx
Oct 2, 2026
Summary
User reports Qwen3-Next Flash IQ2_XS at around 28 tok/s for short generation on a cluster of 4 BC-250 boards, dropping to 24 tok/s at 50k context with around 115 t/s prefill.
Setup is llama.cpp with Vulkan and RPC over 1Gb Ethernet, 100k context, across 6 BC-250 ex-mining boards in an ASRock 4U12G case.
The other two boards run Qwen3.6 35B Q4 at 60 tok/s with 100k context and 450 t/s prefill.