llamaperf

Qwen3.6 35B (3B active)

on NVIDIA RTX 4060 Ti 8GB · llama.cpp · 131,072 ctx

Tone: positive
Oct 6, 2026
Throughput
52-65 t/s gen
Quant
Q4_K_XL (GGUF)
System RAM
64 GB
VRAM reported
8 GB

Use cases

agentic

Summary

User reports Qwen3.6-35B-A3B at 52-65 tok/s on an RTX 4060 Ti 8GB with 64 GB system RAM. Setup is llama.cpp with Q4_K_XL at 131k context, experts offloaded to system RAM and the rest in VRAM, on headless Linux. The user compares download defaults (~25 tok/s), tuned Windows (39-45 tok/s), and tuned headless Linux (52-65 tok/s). They also report Qwen3.8-Flash-Next 125B at 17-19 tok/s and Ternary Bonsai 27B at 36 tok/s.