llamaperf
Oct 6, 2026
Throughput
46.0 t/s gen · 1400.0 t/s pp
System RAM
128 GB

Use cases

coding

Summary

User reports Qwen3.8-Flash-Next at 1,400 t/s prefill and 46 t/s decode on a Strix Halo 128GB laptop at 70 W. Setup is Gufo as the engine, with the model running at xhigh effort. The user notes the model ran on an older Halogen version (0.14.0) and Halogen dropped the connection once, so the last part ran on Gufo. The user compares the local model against Claude Opus 5.5 on a coding task, finding the local model's PR better in tests and edge cases, though Opus was 2 to 10 times faster overall.