Strata
An inference engine for running open-weight LLMs locally.
24 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running Strata
| GPU | VRAM | Reports | Median t/s |
|---|---|---|---|
| NVIDIA RTX 3090nvidia | 24GB | 9 | no plain run of this model |
| NVIDIA RTX 4090nvidia | 24GB | 3 | no plain run of this model |
| NVIDIA RTX 4070 Ti Supernvidia | 16GB | 2 | no plain run of this model |
| NVIDIA RTX 4080 Supernvidia | 16GB | 2 | no plain run of this model |
| NVIDIA RTX 5060 Ti 16GBnvidia | 16GB | 2 | no plain run of this model |
| NVIDIA RTX 5070nvidia | 12GB | 2 | no plain run of this model |
| NVIDIA RTX 5080nvidia | 16GB | 2 | no plain run of this model |
| NVIDIA RTX 5090nvidia | 32GB | 2 | no plain run of this model |
Strata against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
No matched pair yet. No card has plain runs of one model size at one bit class on Strata and on another engine, so llamaperf can't say how it compares on speed. Add a run.
Strata results by GPU
Every card people have run Strata on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
Strata on NVIDIA RTX 30908 reports
- Qwen3.8 125B · 6B active · Q6_K_XL43.9 t/s, 1,817 t/s prompt(2 cards, speculative, part in system RAM)
- Qwen3.8 125B · 6B active · UD-Q4_K_XL97.0 t/s(2 cards, speculative, part in system RAM)
- GLM-5.3 320B · 18B active · UD-Q4_K_XL19.8 t/s, 438 t/s prompt(part in system RAM)
- Qwen3.8 125B · 6B active · IQ3_S116.0 t/s, 2,908 t/s prompt(2 cards, speculative)
- Qwen3.8 125B · 6B active · UD-Q4_K_XL1,700 t/s prompt(part in system RAM)
Strata on NVIDIA RTX 40903 reports
- Qwen3.8 125B · 6B active · IQ3_S120.1 t/s(speculative, part in system RAM)
- Qwen3.8 125B · 6B active · IQ2_XS106.1 t/s, 147 t/s prompt(speculative, part in system RAM)
- Qwen3.8 125B · 6B active · IQ2_XS166.4 t/s(speculative, part in system RAM)
Strata on NVIDIA RTX 4070 Ti Super2 reports
- Qwen3.8 · Q2197.0 t/s, 3,300 t/s prompt(part in system RAM)
- Qwen3.8 125B · 6B active · IQ3_S72.0 t/s, 1,700 t/s prompt(part in system RAM)
Strata on NVIDIA RTX 4080 Super2 reports
- Qwen3.8 125B · 6B active · q3_xss1,000 t/s prompt(part in system RAM)
- Qwen3.8 125B · 6B active · IQ3_S40.0 t/s(part in system RAM)
Strata on NVIDIA RTX 50702 reports
- Qwen3.8 125B · 6B active · IQ3_S53.0 t/s, 1,620 t/s prompt(speculative, part in system RAM)
- Qwen3.8 125B · 6B active · IQ3_XXS44.8 t/s, 414 t/s prompt(part in system RAM)
Strata on NVIDIA RTX 50802 reports
- Qwen3.8 125B · 6B active · UD-Q4_K_XL110.6 t/s, 2,773 t/s prompt(2 cards, speculative, part in system RAM)
- Qwen3.8 125B · 6B active · IQ2_XS1,940 t/s prompt(2 cards, speculative, part in system RAM)
Strata on NVIDIA V100 16GB2 reports
- Qwen3.8 125B · 6B active · Q2_052.0 t/s, 399 t/s prompt(speculative, part in system RAM)
- Qwen3.8 125B · 6B active · UD-Q4_K_XL7,090 t/s prompt(4 cards, speculative, part in system RAM)
Strata on AMD RX 6900 XT 16GB1 report
- Qwen3.8 125B · 6B active · GSQ-RCO25.0 t/s(part in system RAM)
Strata on AMD RX 7600 8GB1 report
- Qwen3.8 125B · 6B active · IQ2_XS24.0 t/s, 68 t/s prompt(part in system RAM)
Strata on AMD Strix Halo 128GB1 report
- Qwen3.8 125B · 6B active · UD-IQ4_XS53.8 t/s, 1,293 t/s prompt
Strata on NVIDIA RTX 3060 12GB1 report
- Qwen3.8 125B · 6B active27.9 t/s(part in system RAM)
Strata on NVIDIA RTX 3080 20GB1 report
- Qwen3.8 125B · 6B active · IQ3_XXS105.0 t/s, 5,000 t/s prompt(4 cards)
Strata on NVIDIA RTX 3090 Ti1 report
- Qwen3.8 125B · 6B active · IQ3-XXS105.0 t/s, 2,000 t/s prompt(2 cards, speculative)
Strata on NVIDIA RTX 40801 report
- Qwen3.8 125B · 6B active · Q355.0 t/s, 590 t/s prompt(part in system RAM)
Strata on NVIDIA RTX 5060 Ti 16GB1 report
- Qwen3.8 · IQ1_M55.0 t/s, 1,500 t/s prompt(part in system RAM)
Strata on NVIDIA RTX 5070 Ti1 report
- Qwen3.8 125B · 6B active · IQ3_S53.5 t/s(speculative, part in system RAM)
Strata on NVIDIA RTX 5070 Ti Laptop 12GB1 report
- Qwen3.8 125B · 6B active · IQ3_XXS51.0 t/s, 1,500 t/s prompt(part in system RAM)
Strata on NVIDIA RTX 50901 report
- Qwen3.8 125B · 6B active · IQ3_XXS1,796 t/s prompt(2 cards, speculative, part in system RAM)
Strata on NVIDIA RTX Pro 4500 Blackwell 32GB1 report
- Qwen3.8 125B · 6B active · NVFP41,680 t/s prompt(speculative, part in system RAM)
Strata, GPU not identified6 reports
- Qwen3.8 125B · 6B active · IQ3_XXS104.0 t/s(4 cards, part in system RAM)
- Qwen3.8 125B · 6B active · IQ2_XS800 t/s prompt
- Qwen3.8 125B · 6B active · IQ3_XXS58.0 t/s, 1,563 t/s prompt
- Qwen3.8 125B · 6B active · IQ2_XS85.0 t/s(part in system RAM)
- Qwen3.8 125B · 6B active · q2_010.0 t/s(speculative, part in system RAM)
Top models on Strata
Frequently asked
Is Strata faster than other engines?
llamaperf has no matched comparison for Strata yet: no card has plain runs of the same model size at the same bit class on Strata and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.