llamaperf

Qwen3.8 27B Escha-W2

on NVIDIA RTX 4090 · SGLang · 65,536 ctx

Oct 3, 2026
Throughput
67.0 t/s gen · 2600-2820 t/s pp
Quant
escha 2-bit (safetensors)
VRAM reported
24 GB

Use cases

codingmath

Summary

User reports Qwen3.8-27B-Escha-W2 at 67.0 tok/s single-stream decode on an RTX 4090. Setup is SGLang with a 2-bit escha quant (2.469 bits/weight, 10.15 GB of weights) at 64k context, CUDA graphs on, prefix caching off. With the model's own MTP head as a speculative draft the same card reaches 129.3 tok/s (1.92x). Peak server throughput is 649 tok/s at 16 streams; prefill is ~2,600-2,820 tok/s and TTFT at a 2k prompt is 0.73 s. The card also ran 87.1 tok/s on an RTX 5090 and 40.7 tok/s on an RTX 3090 with ESCHA_ROUTE=blackwell.