llamaperf

gpt-oss

OpenAI · 2 reports

As of 11 Oct 2026, gpt-oss 117B · 5.1B active at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

gpt-oss VRAM requirements by size and quant →

How does gpt-oss run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for gpt-oss on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run gpt-oss yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for gpt-oss

Filter this model’s reports by setup →
Thin page (2 of 3 reports needed for indexing). Add yours.
AMD hardware
generation:
54.1 tokens/s
prompt processing (prefill):
645.1 tokens/s
quant:
MXFP4 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.

agentictool-use

User reports gpt-oss-120b at 54.11 t/s generation and 645.07 t/s prompt processing on a GMKtec EVO-X2 with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB unified memory. Setup is llama.cpp build 11456 with MXFP4 GGUF weights, Vulkan (Mesa RADV) backend, flash attention on, whole model on the iGPU, no speculative decoding, one request at a time. Three runs after a warm-up gave 54.11, 54.64 and 54.50 t/s generation with 645.07, 656.81 and 657.14 t/s prompt processing on an 839-token prompt. Two small embedding and reranker models were loaded but idle. The user runs this model for agentic tool use and research.

Oct 9, 2026
Tone: positiveNVIDIA hardware
tool-use

User reports gpt-oss-20b at close to 200 tps aggregate across parallel requests with the full 128k context on a single RTX 3090, under a custom harness (burrito) that fixes the model's tool calling; 320,192 evals over 8 seeds and 3.49B tokens took 1,062 GPU hours at batch size 1. Setup is factory-precision gpt-oss-20b weights on one RTX 3090 serving parallel requests; no single-stream figure and no decoding method is given. Effort levels affect accuracy: Low 38.3%, Medium 97.1%, High 100.0%. The user mentions llama.cpp and vLLM had issues with the model.

Sep 7, 2026

Get a weekly email of new gpt-oss reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
gpt-oss 117B (5.1B active)
AMD Strix Halo 128GB
MXFP4
llama.cpp
Not reported54.1 tokens/s