llamaperf
Sep 18, 2026
Throughput
600.0 t/s gen

Use cases

codingsummarization

Summary

User reports Qwen3.6 35B-A3B at 600 t/s single request on an RTX Pro 6000 Blackwell. Setup uses the Ninfer engine. User describes it as a drudgework model for read-and-find or code tasks, noting it may use 20x more tokens but is still faster than many local models, and compares it favorably to Cerebras.