llamaperf

Qwen3.5 9B

on NVIDIA V100 16GB · ik_llama.cpp

Sep 26, 2026
Throughput
4.6 t/s gen · 53.0 t/s pp
Quant
Q6_K (GGUF)
System RAM
256 GB
VRAM reported
16 GB

Summary

User asks for help building an offline LLM rig and reports a test of Qwen3.5 9B at Q6_K on their main rig, getting 53 t/s prompt processing and 4.6 t/s generation with ik_llama.cpp. Setup is ik_llama.cpp with Q6_K; the test rig's hardware is not specified. The planned build is dual Xeon 2683v4 with 256GB DDR4 ECC and a V100 16GB, targeting DeepSeek V4 Flash and MiMo V2.5 at UD-IQ4_XS. User estimates 6-7 t/s generation and 50-60 t/s prompt processing for DeepSeek V4 Flash on CPU alone, hoping a GPU raises prompt processing to 200-300 t/s, and expects MiMo V2.5 to be 1-2 t/s slower. ik_llama.cpp doubled prompt processing and added 1 t/s generation over llama.cpp in their test.