llamaperf

Nex-N2.5-mini

1 report

Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: positive
reported speed:
76.9 tokens/s generation
quant:
ROCmFP4 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post benchmarks Nex-N2.5-mini (ROCmFP4 GGUF) on AMD Strix Halo against Qwen3.8-27B. Decode 76.9 tok/s vs Qwen3.8-27B's 14-34 tok/s. Terminal-Bench 2.1: 73.4 vs 73.0; WebArena: 63.4 vs 64.8; SWE-Bench: 43.8 vs 61.7. Weights at huggingface.co/julianmb/Nex-N2.5-mini-ROCmFP4-GGUF; engine HaloFPX (github.com/julianmb/halofpx).