llamaperf

Swift-1.5

1 report

Swift-1.5 VRAM requirements by size and quant →

How does Swift-1.5 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Swift-1.5 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Swift-1.5 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Swift-1.5

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: positive
reported speed:
112.0 tokens/s generation
quant:
MXFP4 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports 112 tok/s decode on coding prompts with WHIRL, the engine they wrote, for Swift-1.5 27B MXFP4 on a single Radeon AI PRO R9700, against 61 tok/s for llama.cpp b11214 on the same GGUF. Setup is one Radeon AI PRO R9700 32GB over USB4 on Windows 11, MXFP4 weights, WHIRL v0.1.3 with speculative decoding, single request. Prefill reaches 1,757 tok/s at 128K and 970 at 256K, and 4 concurrent users get 182 tok/s in total. On Ornith-1.5-35B-A3B MXFP4 WHIRL decodes at 258 tok/s and prefills over 11,000 tok/s at 8K.

Oct 8, 2026

Get a weekly email of new Swift-1.5 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Swift-1.5 27B
AMD Radeon AI PRO R9700 32GB
MXFP4
WHIRL
Not reported112.0 tokens/s