llamaperf

Qwen3.5 9B DeepSeek-V4-Flash

on Intel Arc A770 16GB · llama.cpp · 262,144 ctx

Tone: positive
Oct 7, 2026
Throughput
49.0 t/s gen · 48.0 t/s pp
Quant
Q6_K (GGUF)
KV cache
q4_1
Flash Attention
on
System RAM
32 GB
VRAM reported
16 GB

Use cases

codingtool-usevisionlong-contextagentic

Summary

User reports Qwen3.5-9B at 49 t/s generation and 48 t/s prompt processing on an Intel Arc A770 16GB. Setup is llama.cpp b9521 (Vulkan) with Q6_K weights, q4_1 KV cache, 256K context, flash attention, vision mmproj and MTP speculative decoding on a single slot. The same model under WSL2 with Q8_0 KV cache and 128K context reached about 40 t/s without vision. Qwopus3.5-4B-Coder hit 64 t/s generation and 100 t/s prompt at 96K context, and Gemma 4 12B reached 22 t/s generation and 74 t/s prompt at 128K context.