llamaperf

Cyber-Tiel 35B (3B active) Coder

on NVIDIA RTX 4060 Laptop 8GB · llama.cpp · 131,072 ctx

Sep 23, 2026
Throughput
23.0 t/s gen · 35.0 t/s pp
Quant
UD-IQ3_XXS (GGUF)
KV cache
q4_0
System RAM
16 GB
VRAM reported
8 GB

Use cases

coding

Summary

User reports Cyber-Tiel Coder 35B-A3B at 23 t/s decode and 35 t/s prefill on an RTX 4060 Laptop 8GB with 16 GB system RAM, at 131,072 context. Setup is llama.cpp with the UD-IQ3_XXS GGUF, q4_0 KV cache, flash attention, 30 MoE layers offloaded to CPU, and MTP speculative decoding with 2 draft tokens. Figures are at 4K context and depend on free RAM for mmap caching; a 13-prompt benchmark table is promised later.