llamaperf

Qwen3.8 125B (6B active) Flash-Next

on 4× NVIDIA RTX 3080 20GB · Strata · 150,000 ctx

Tone: positive
Oct 5, 2026
Throughput
105.0 t/s gen · 5000.0 t/s pp
Quant
IQ3_XXS (GGUF)

Use cases

agenticlong-context

Summary

User reports Qwen3.8 Flash-Next 125B at 105 t/s generation and 5000 t/s prompt processing on 4x RTX 3080 20GB modded cards. Setup is Strata with IQ3_XXS quant at 150000 context length, running via docker-compose on Unraid. User says they are absolutely impressed and tested with a Hermes agent at 150k context.