Qwen3.5 9B
on NVIDIA V100 16GB · ik_llama.cpp
Sep 26, 2026
Summary
User asks for help building an offline LLM rig and reports a test of Qwen3.5 9B at Q6_K on their main rig, getting 53 t/s prompt processing and 4.6 t/s generation with ik_llama.cpp.
Setup is ik_llama.cpp with Q6_K; the test rig's hardware is not specified. The planned build is dual Xeon 2683v4 with 256GB DDR4 ECC and a V100 16GB, targeting DeepSeek V4 Flash and MiMo V2.5 at UD-IQ4_XS.
User estimates 6-7 t/s generation and 50-60 t/s prompt processing for DeepSeek V4 Flash on CPU alone, hoping a GPU raises prompt processing to 200-300 t/s, and expects MiMo V2.5 to be 1-2 t/s slower. ik_llama.cpp doubled prompt processing and added 1 t/s generation over llama.cpp in their test.