llamaperf

Qwen3.8 Flash-Next Coder

on NVIDIA RTX 5060 Ti 16GB · Strata · 65,000 ctx

Tone: positive
Oct 6, 2026
Throughput
55.0 t/s gen · 1500.0 t/s pp
Quant
IQ1_M (GGUF)
System RAM
32 GB
VRAM reported
16 GB

Use cases

coding

Summary

User reports Qwen3.8 Flash-Next Coder at 55 tok/s generation and 1500 tok/s prompt processing on an RTX 5060 Ti 16GB with 32GB system RAM. Setup is Strata with IQ1_M quant at 65k context. A one-shot landing page demo ran at 46 tok/s over about 2 minutes; an interactive 3D globe took 8 minutes plus a minute of JS fixes. User says the IQ1_M output quality surprised them, comparable to above Q3 for other local models.