FreeToken
An inference engine for running open-weight LLMs locally.
9 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running FreeToken
| GPU | VRAM | Reports | Median t/s, Qwen3.8 27B 2-bit |
|---|---|---|---|
| NVIDIA RTX 5090nvidia | 32GB | 4 | no plain run of this model |
| NVIDIA RTX 3090nvidia | 24GB | 2 | no plain run of this model |
| Intel Arc B580 12GBintel | 12GB | 1 | 0.7 |
| NVIDIA RTX 3060 12GBnvidia | 12GB | 1 | no plain run of this model |
| NVIDIA RTX 4060 Laptop 8GBnvidia | 8GB | 1 | no plain run of this model |
FreeToken against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
No matched pair yet. No card has plain runs of one model size at one bit class on FreeToken and on another engine, so llamaperf can't say how it compares on speed. Add a run.
FreeToken results by GPU
Every card people have run FreeToken on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
FreeToken on NVIDIA RTX 50904 reports
- Qwen3.8 125B · 6B active · NVFP426.2 t/s(part in system RAM)
- Qwen3.8 125B · 6B active · NVFP450.5 t/s(part in system RAM)
- Qwen3.8 125B · 6B active · NVFP450.0 t/s, 2,300 t/s prompt(part in system RAM)
- Qwen3.8 125B · 6B active73.6 t/s
FreeToken on NVIDIA RTX 30902 reports
- DeepSeek V4 Flash · fp85.6 t/s(2 cards, part in system RAM)
- Qwen3.8 125B · 6B active · NVFP448.0 t/s(2 cards, part in system RAM)
FreeToken on Intel Arc B580 12GB1 report
- Qwen3.8 27B · UD-Q2_K_XL0.7 t/s
FreeToken on NVIDIA RTX 4060 Laptop 8GB1 report
- Qwen3.6 35B · 3B active39.0 t/s(part in system RAM)
FreeToken, GPU not identified1 report
- Qwen3.820.0 t/s(part in system RAM)
Top models on FreeToken
Frequently asked
Is FreeToken faster than other engines?
llamaperf has no matched comparison for FreeToken yet: no card has plain runs of the same model size at the same bit class on FreeToken and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.