Qwen3.8 27B Escha-W2
RTX 4080 Super · SGLang · 98,304 ctx
- reported speed:
- 59.0 tokens/s generation
- quant:
- 2.469 bpw
- kv:
- FP8 E4M3
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingagenticlong-context
User pushed Qwen3.8-27B-Escha-W2 to 98K context on a 16GB 4080 Super. Reports ~59 tok/s short context, ~50.5 tok/s at 60K context. Quality surprisingly good despite aggressive 2.469 bpw quant. MTP4 gave ~67.7 tok/s at 64K context but chose no speculation for max context. Setup uses Escha's SGLang build with FP8 KV cache and BF16 SSM state.