llamaperf

Qwen3.8 35B

on RTX 5050 8GB · FreeToken

Oct 4, 2026
Throughput
20.0 t/s gen
System RAM
32 GB
VRAM reported
8 GB

Summary

User reports Qwen3.8 35B at about 20 tokens per second on an RTX 5050 8GB with 32 GB of DDR4 RAM. Setup is FreeToken; the model runs with system RAM involved since 35B does not fit in 8 GB of VRAM. User asks whether switching to Strata would help, and states hardware upgrades are not an option.