Cyber-Tiel 35B (3B active) Coder
NVIDIA RTX 4060 Laptop 8GB · llama.cpp · 131,072 ctx
- reported speed:
- 23.0 tokens/s generation · 35.0 tokens/s prompt processing
- quant:
- UD-IQ3_XXS (GGUF)
- kv:
- q4_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Cyber-Tiel Coder 35B-A3B at 23 t/s decode and 35 t/s prefill on an RTX 4060 Laptop 8GB with 16 GB system RAM, at 131,072 context. Setup is llama.cpp with the UD-IQ3_XXS GGUF, q4_0 KV cache, flash attention, 30 MoE layers offloaded to CPU, and MTP speculative decoding with 2 draft tokens. Figures are at 4K context and depend on free RAM for mmap caching; a 13-prompt benchmark table is promised later.