llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 5070 Ti · Strata · 262,144 ctx

Tone: positive
Oct 4, 2026
Throughput
53.5 t/s gen
Quant
IQ3_S (GGUF)
System RAM
96 GB
VRAM reported
16 GB

Use cases

agenticcodinglong-context

Summary

User reports Qwen3.8-Flash-Next at 53.5 tok/s with a cached prefix on an RTX 5070 Ti 16 GB, up from 17.2 tok/s before tuning. Setup is Strata with IQ3_S at 262K context, KV cache in RAM streaming to a 32K VRAM window, speculative decoding with depth 6 and min-p 0.70, and 13 pool workers on a 14700KF with 96 GB RAM. Cold run was 43 tok/s; the 53.5 figure is with the prefix cached, which the user says is what agent sessions actually do. The user also reports the calibrator's own curve at 42.1 tok/s at --pcie-frac 0.0 down to 20.7 at 0.75.