Qwen3.8 27B
Instinct MI300X 192GB · SGLang · 1,000,000 ctx
User reports Qwen3.8-27B at 495 t/s on a single MI300X at 1M context, up from 311 t/s. Setup is dstack's toolkit with source-level patches to SGLang's AITER attention backend. The figure is for four concurrent users at 10k in / 1.5k out, with p50 TTFT under 1.5s.