Qwen3.8 125B (6B active) Flash-Next
on 4× NVIDIA RTX 3080 20GB · Strata · 150,000 ctx
Oct 5, 2026
Use cases
agenticlong-context
Summary
User reports Qwen3.8 Flash-Next 125B at 105 t/s generation and 5000 t/s prompt processing on 4x RTX 3080 20GB modded cards.
Setup is Strata with IQ3_XXS quant at 150000 context length, running via docker-compose on Unraid.
User says they are absolutely impressed and tested with a Hermes agent at 150k context.