Qwen3.6 35B (3B active)
on AMD RX 6600 XT · llama.cpp · 65,536 ctx
Sep 23, 2026
Use cases
agentic
Summary
User asks what performance P100 owners get, having ordered one for $80, and reports their current baseline on an RX 6600 XT with 32GB DDR4 3600.
Current setup runs unsloth Qwen3.6 35B-A3B UD_Q4_K_XL in llama.cpp with MTP and --cpu-moe at 64k full-precision context, giving 30 t/s generation and 48-50 t/s decode, with prefill up to 800 t/s at 0 context and 700 t/s at 10k.
User hopes the P100 can match the decode numbers and plans to tune -b and -ub for its higher core count; the card will go into a dedicated inference machine.