llamaperf

Qwen3.6 35B (3B active)

on AMD RX 6600 XT · llama.cpp · 65,536 ctx

Sep 23, 2026
Throughput
30.0 t/s gen
Quant
UD_Q4_K_XL (GGUF)
MTP (Multi-Token Prediction)
on
System RAM
32 GB
VRAM reported
16 GB

Use cases

agentic

Summary

User asks what performance P100 owners get, having ordered one for $80, and reports their current baseline on an RX 6600 XT with 32GB DDR4 3600. Current setup runs unsloth Qwen3.6 35B-A3B UD_Q4_K_XL in llama.cpp with MTP and --cpu-moe at 64k full-precision context, giving 30 t/s generation and 48-50 t/s decode, with prefill up to 800 t/s at 0 context and 700 t/s at 10k. User hopes the P100 can match the decode numbers and plans to tune -b and -ub for its higher core count; the card will go into a dedicated inference machine.