- reported speed:
- 46.0 tokens/s generation · 1400.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-Flash-Next at 1,400 t/s prefill and 46 t/s decode on a Strix Halo 128GB laptop at 70 W. Setup is Gufo as the engine, with the model running at xhigh effort. The user notes the model ran on an older Halogen version (0.14.0) and Halogen dropped the connection once, so the last part ran on Gufo. The user compares the local model against Claude Opus 5.5 on a coding task, finding the local model's PR better in tests and edge cases, though Opus was 2 to 10 times faster overall.