llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Looking for a particular model?

Search for a model to see its reported speeds across GPUs and Macs.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Ornith 1.0
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

AMD hardware
generation:
74.1 tokens/s
prompt processing (prefill):
1102.5 tokens/s
quant:
Q4_K_M (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.

Source-attributed historical benchmark by axjns, measured August 17, 2026. This is a public third-party llama-bench measurement, not my own hardware or an independent hardware rerun. The model is Ornith 1.0, NOT Ornith 1.5. RESULT AND ACTUAL WORK The submitted rounded fields are prefill 1102.5 tokens/s and generation 74.1 tokens/s. Exact source means and standard deviations across three repetitions are: - Prefill: 1102.529259 +/- 24.069094 tokens/s, n_prompt=512, n_gen=0, depth=0. - Decode: 74.104028 +/- 0.245648 tokens/s, n_prompt=0, n_gen=128, depth=0. These are TWO SEPARATE synthetic microbenchmark phases. They are not one 512-input/128-output request, not medians, not client end-to-end throughput or TTFT, and not an application conversation. The context field is left blank rather than presenting an input count as configured context capacity. Depth zero means no prefilled context depth for these headline points; it does not establish cold model loading or a cold filesystem cache. RAW SAMPLES AND CHECK The pp512 sample durations are 472523394, 467817691, and 453257747 ns. The tg128 durations are 1728353257, 1732444252, and 1721144960 ns. Arithmetic means of 512e9/ns and 128e9/ns respectively reproduce 1102.529259240522 and 74.10402815760727 tokens/s, agreeing with the published means. The reported standard deviations are the source values. The two records are lines 41 and 42 of the JSONL. Source timestamps for those records are 2026-08-17T13:09:34Z and 2026-08-17T13:09:42Z. Other separate points from the same model, lines 43 and 44, are pp4096 at depth 0: 1060.246020 +/- 8.507957 tokens/s, and tg128 at depth 4096: 69.904247 +/- 0.369516 tokens/s. These are supplementary observations, not pooled into the headline rates. These additional 4K-scale observations do not establish long-context application performance. HARDWARE AND MEMORY One AMD Ryzen AI MAX+ 395 / Radeon 8060S Graphics, gfx1151, with 128 GB unified memory according to the dataset card. GPU count 1 denotes one machine/iGPU. The raw records expose a 64 GiB GPU allocation (68719476736 bytes) and 62 GiB OS-visible system RAM. The 64 GiB carve-out and 62 GiB OS-visible figure are not extra memory to add to the 128 GB total. The separate reported-VRAM form field is deliberately blank to avoid implying independent dedicated VRAM. For the headline prefill record, source GPU-allocation peak is 23.66 GiB and baseline 3.60 GiB, difference 20.06 GiB. For decode they are 23.15, 3.59, and 19.56 GiB. These are rocm-smi readings sampled at 4 Hz; a sampled maximum is not a guaranteed instantaneous peak, process RSS or total resident system-memory measurement. All four Ornith rows have gpu_contended=false and an empty concurrent_workloads list. Power profile, watts, clock rates and temperature are unreported. RUNTIME AND BACKEND CONFLICT Use the raw runtime provenance: llama.cpp llama-bench build 9590, recorded commit d2462f8f7; lb_backends=BLAS,Vulkan and lb_gpu_info=AMD Radeon 8060S Graphics (RADV GFX1151). The dataset card describes ROCm, and the environment contains ROCm/HIP 7.13.26162, but those installed versions do NOT make this a HIP/ROCm inference measurement. The actual recorded benchmark backend is Vulkan/RADV. Mesa version and an independently verified binary hash are unknown. Fedora release 43, kernel 7.0.14-101.fc43.x86_64. n_batch=2048, n_ubatch=512, n_threads=16, n_gpu_layers=999, n_cpu_moe=0, split_mode=layer, main_gpu=0, no_kv_offload=false, use_mmap=true, use_direct_io=false. K and V cache types are both f16. lb_flash_attn=-1 records AUTO; actual Flash Attention engagement is unverified and the form flag is left unknown. No speculative route or draft model is recorded. No inference-quality or speculation-quality claim is made. MODEL IDENTITY Recorded filename: ornith-1.0-35b-Q4_K_M.gguf. Raw model type: qwen35moe 35B.A3B Q4_K - Medium. Raw parameter count is 34660610688; the form uses nominal 35B total / 3B active. The GGUF file is 21166757760 bytes on disk; lb_model_size is 21155768832 bytes. The source parses the quant label from the filename, so Q4_K_M is filename-derived rather than independently verified by inspecting the weights. The tested model Hugging Face repository, tested model revision, and tested weight-file hash are UNREPORTED. Do not infer a model download identity or immutable weight pin from the local folder name or from a current similarly named model page. The Hugging Face links below point to a BENCHMARK DATASET and its harness, not to the tested model weights. The dataset commit is an evidence snapshot, not the model revision. PUBLIC PRIMARY SOURCES Dataset card and methodology: https://huggingface.co/datasets/axjns/strix-halo-inference-bench Immutable raw benchmark records (lines 41-44): https://huggingface.co/datasets/axjns/strix-halo-inference-bench/blob/9eecbcec0c626e9e8340c47752ebee50112c856a/data/results.jsonl Current raw-record view: https://huggingface.co/datasets/axjns/strix-halo-inference-bench/blob/main/data/results.jsonl Published measurement harness: https://huggingface.co/datasets/axjns/strix-halo-inference-bench/blob/main/strix_bench.py Verification October 11, 2026: all 44 records were readable in the primary file view; pinned and main snapshots parsed identically. Headline means were independently recomputed from published timing samples. Live catalog dedup checked all 115 Strix Halo reports across six pages. The existing Ornith 1.5 ROCmFP4 report is a different model/version and setup. These checks verify evidence consistency, not independent hardware reproduction. This single-machine, single-build snapshot measures speed only; model quality, real-task usefulness, longer contexts and sustained service behavior were not evaluated.

Oct 11, 2026
Showing 1–1 of 1
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090185NVIDIA RTX 5090131AMD Strix Halo 128GB120NVIDIA DGX Spark72NVIDIA RTX Pro 6000 Blackwell59NVIDIA RTX 5060 Ti 16GB57NVIDIA RTX 3060 12GB50AMD Radeon AI PRO R9700 32GB50NVIDIA RTX 409042NVIDIA RTX 5070 Ti34

Records by model

1670 total
Qwen3.8922
Qwen3.6193
DeepSeek V4 Flash138
Gemma 473
Qwen3.546
GLM-5.334
other264

Records by engine

1271 total
llama.cpp646
vLLM182
Strata66
NInfer44
oMLX40
other293

Use cases

coding 531agentic 346long-context 256tool-use 148vision 103summarization 52math 42creative-writing 36multilingual 21text-generation 9rp 8reasoning 3
coding531agentic346long-context256tool-use148vision103summarization52math42creative-writing36

Median generation t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509082RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX48RTX 309036RTX 5090 Laptop 24GB34RTX 409032V100 32GB31Radeon AI PRO R9700 32GB30RX 7800 XT 16GB30

Prompt processing, same model

M5 Ultra 256GB1800RX 7900 XTX530RTX 3090561RTX 40901949V100 32GB941RTX 3080 20GB997Arc Pro B70512M1 Ultra 128GB153Arc Pro B60 24GB380M1 Max 32GB82

Reports by model size

Qwen3.8 27B535Qwen3.8 125B · 6B active355DeepSeek V4 Flash 284B · 13B active129Qwen3.6 35B · 3B active120Qwen3.6 27B71GLM-5.3 320B · 18B active30Gemma 4 26B · 4B active30DeepSeek V4.1 Flash 552B · 16B active27

Quants

Q4_K_M144NVFP4116Q4_K_XL66IQ4_XS63UD-Q4_K_XL62IQ3_XXS44Q442Q8_039MXFP4334bit32