Qwen3.5 35B (3B active)
NVIDIA RTX 4060 Ti 8GB · llama.cpp
- reported speed:
- 8.0 tokens/s generation
- quant:
- Q2 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 8 t/s with Qwen3.5 35B Q2 on an RTX 4060 Ti 8GB, with no CUDA device selected. Setup is llama.cpp with a Q2 GGUF quant; the user notes VRAM usage was 3000MB/8k and that the run had no CUDA device selected. The user is troubleshooting a cuBLAS crash and asks how to cleanly uninstall and reinstall llama.cpp with CUDA 12, and whether CUDA would improve generation speed.