llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 4070 Ti Super · Strata · 262,144 ctx

Tone: positive
Oct 5, 2026
Throughput
72.0 t/s gen · 1700.0 t/s pp
Quant
IQ3_S (GGUF)
System RAM
128 GB
VRAM reported
16 GB

Use cases

codingtool-uselong-context

Summary

User reports Qwen3.8 Flash-Next at 72 t/s generation and ~1,700 t/s prompt reading on an RTX 4070 Ti Super 16GB. Setup is Strata 0.1.39 with IQ3_S (~84GB) at 256K context, on a Ryzen 9 7950X3D with 128GB DDR5. Out of the box it did 8 t/s; after START-HERE.bat --calibrate it reached 72 t/s. A 17K prompt reads in 8 seconds. User's 27-task coding suite scored 27/27 x3 and HumanEval 159/164. Same model family under llama.cpp with offloading ran at 19 t/s. Keeping the 29GB n-gram table in RAM with --ple-io mmap made results worse (25-26/27 and slower).