llamaperf
Oct 7, 2026
Throughput
800.0 t/s pp
Quant
IQ2_XS (GGUF)
System RAM
32 GB
VRAM reported
16 GB

Use cases

agenticcoding

Summary

User asks whether Swift 1.5 Flash-Next at IQ3_XXS will fit in 16GB VRAM and 32GB RAM, having run the IQ2_XS build on Strata at an average 50 t/s and 800 t/s prompt processing, sometimes up to 73 t/s. Setup is Strata with an IQ2_XS GGUF of the Flash-Next model. The user wants a higher quantization for lower KLD and better agentic coding accuracy and asks for others' configs.