Ornith 1.0 35B (3B active)
AMD Strix Halo 128GB · llama.cpp
- generation:
- 74.1 tokens/s
- prompt processing (prefill):
- 1102.5 tokens/s
- quant:
- Q4_K_M (gguf)
- kv:
- F16
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.
Source-attributed historical benchmark by axjns, measured August 17, 2026. This is a public third-party llama-bench measurement, not my own hardware or an independent hardware rerun. The model is Ornith 1.0, NOT Ornith 1.5. RESULT AND ACTUAL WORK The submitted rounded fields are prefill 1102.5 tokens/s and generation 74.1 tokens/s. Exact source means and standard deviations across three repetitions are: - Prefill: 1102.529259 +/- 24.069094 tokens/s, n_prompt=512, n_gen=0, depth=0. - Decode: 74.104028 +/- 0.245648 tokens/s, n_prompt=0, n_gen=128, depth=0. These are TWO SEPARATE synthetic microbenchmark phases. They are not one 512-input/128-output request, not medians, not client end-to-end throughput or TTFT, and not an application conversation. The context field is left blank rather than presenting an input count as configured context capacity. Depth zero means no prefilled context depth for these headline points; it does not establish cold model loading or a cold filesystem cache. RAW SAMPLES AND CHECK The pp512 sample durations are 472523394, 467817691, and 453257747 ns. The tg128 durations are 1728353257, 1732444252, and 1721144960 ns. Arithmetic means of 512e9/ns and 128e9/ns respectively reproduce 1102.529259240522 and 74.10402815760727 tokens/s, agreeing with the published means. The reported standard deviations are the source values. The two records are lines 41 and 42 of the JSONL. Source timestamps for those records are 2026-08-17T13:09:34Z and 2026-08-17T13:09:42Z. Other separate points from the same model, lines 43 and 44, are pp4096 at depth 0: 1060.246020 +/- 8.507957 tokens/s, and tg128 at depth 4096: 69.904247 +/- 0.369516 tokens/s. These are supplementary observations, not pooled into the headline rates. These additional 4K-scale observations do not establish long-context application performance. HARDWARE AND MEMORY One AMD Ryzen AI MAX+ 395 / Radeon 8060S Graphics, gfx1151, with 128 GB unified memory according to the dataset card. GPU count 1 denotes one machine/iGPU. The raw records expose a 64 GiB GPU allocation (68719476736 bytes) and 62 GiB OS-visible system RAM. The 64 GiB carve-out and 62 GiB OS-visible figure are not extra memory to add to the 128 GB total. The separate reported-VRAM form field is deliberately blank to avoid implying independent dedicated VRAM. For the headline prefill record, source GPU-allocation peak is 23.66 GiB and baseline 3.60 GiB, difference 20.06 GiB. For decode they are 23.15, 3.59, and 19.56 GiB. These are rocm-smi readings sampled at 4 Hz; a sampled maximum is not a guaranteed instantaneous peak, process RSS or total resident system-memory measurement. All four Ornith rows have gpu_contended=false and an empty concurrent_workloads list. Power profile, watts, clock rates and temperature are unreported. RUNTIME AND BACKEND CONFLICT Use the raw runtime provenance: llama.cpp llama-bench build 9590, recorded commit d2462f8f7; lb_backends=BLAS,Vulkan and lb_gpu_info=AMD Radeon 8060S Graphics (RADV GFX1151). The dataset card describes ROCm, and the environment contains ROCm/HIP 7.13.26162, but those installed versions do NOT make this a HIP/ROCm inference measurement. The actual recorded benchmark backend is Vulkan/RADV. Mesa version and an independently verified binary hash are unknown. Fedora release 43, kernel 7.0.14-101.fc43.x86_64. n_batch=2048, n_ubatch=512, n_threads=16, n_gpu_layers=999, n_cpu_moe=0, split_mode=layer, main_gpu=0, no_kv_offload=false, use_mmap=true, use_direct_io=false. K and V cache types are both f16. lb_flash_attn=-1 records AUTO; actual Flash Attention engagement is unverified and the form flag is left unknown. No speculative route or draft model is recorded. No inference-quality or speculation-quality claim is made. MODEL IDENTITY Recorded filename: ornith-1.0-35b-Q4_K_M.gguf. Raw model type: qwen35moe 35B.A3B Q4_K - Medium. Raw parameter count is 34660610688; the form uses nominal 35B total / 3B active. The GGUF file is 21166757760 bytes on disk; lb_model_size is 21155768832 bytes. The source parses the quant label from the filename, so Q4_K_M is filename-derived rather than independently verified by inspecting the weights. The tested model Hugging Face repository, tested model revision, and tested weight-file hash are UNREPORTED. Do not infer a model download identity or immutable weight pin from the local folder name or from a current similarly named model page. The Hugging Face links below point to a BENCHMARK DATASET and its harness, not to the tested model weights. The dataset commit is an evidence snapshot, not the model revision. PUBLIC PRIMARY SOURCES Dataset card and methodology: https://huggingface.co/datasets/axjns/strix-halo-inference-bench Immutable raw benchmark records (lines 41-44): https://huggingface.co/datasets/axjns/strix-halo-inference-bench/blob/9eecbcec0c626e9e8340c47752ebee50112c856a/data/results.jsonl Current raw-record view: https://huggingface.co/datasets/axjns/strix-halo-inference-bench/blob/main/data/results.jsonl Published measurement harness: https://huggingface.co/datasets/axjns/strix-halo-inference-bench/blob/main/strix_bench.py Verification October 11, 2026: all 44 records were readable in the primary file view; pinned and main snapshots parsed identically. Headline means were independently recomputed from published timing samples. Live catalog dedup checked all 115 Strix Halo reports across six pages. The existing Ornith 1.5 ROCmFP4 report is a different model/version and setup. These checks verify evidence consistency, not independent hardware reproduction. This single-machine, single-build snapshot measures speed only; model quality, real-task usefulness, longer contexts and sustained service behavior were not evaluated.