Qwen3.5 9B DeepSeek-V4-Flash
Intel Arc A770 16GB · llama.cpp · 262,144 ctx
- reported speed:
- 49.0 tokens/s generation · 48.0 tokens/s prompt processing
- quant:
- Q6_K (GGUF)
- kv:
- q4_1
- flash attention:
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.5-9B at 49 t/s generation and 48 t/s prompt processing on an Intel Arc A770 16GB. Setup is llama.cpp b9521 (Vulkan) with Q6_K weights, q4_1 KV cache, 256K context, flash attention, vision mmproj and MTP speculative decoding on a single slot. The same model under WSL2 with Q8_0 KV cache and 128K context reached about 40 t/s without vision. Qwopus3.5-4B-Coder hit 64 t/s generation and 100 t/s prompt at 96K context, and Gemma 4 12B reached 22 t/s generation and 74 t/s prompt at 128K context.