Qwen3.8 Flash-Next
RTX 4070 · Nebula · 24,576 ctx
- reported speed:
- 7.2 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-Flash-Next at 7.19 t/s on an RTX 4070 Ti 12GB with 128GB DDR4 RAM. Setup is the Nebula C/CUDA engine with native MTP speculative decoding, GPU expert caching (27 of 512 experts per layer resident in VRAM), and CPU MoE execution on an Intel i9-9940X using 14 threads, with a 24,576-token context capacity. A limited-tolerance acceptance mode reaches 7.54 t/s. Time to first token is 25.11 seconds for a 512-token input and 113.94 seconds for 2,048 tokens. Prefill is noted as slow for longer prompts.