Qwen3.8 27B
2× Tesla P40 24GB · 150,000 ctx
- reported speed:
- 45.0 tokens/s generation · 450.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at up to 45 t/s generation and 450 t/s prefill on 2x Tesla P40 with fresh context. At 150K+ context prefill falls to around 120 t/s and generation to 12-16 t/s. User notes that running two agents concurrently breaks prefix caching and KV cache sharing, and asks whether a single-supervisor harness with sequential subagent handoffs exists.