llamaperf

Gemma 4 26B

on 2× NVIDIA RTX 4060 · LM Studio

Tone: positiveNVIDIA hardware
Sep 30, 2026
Throughput
75.0 t/s gen · 1500.0 tokens/s prompt processing (prefill)Prompt processing measures input prefill speed; it does not tell you the time to first token.
Quant
IQ (GGUF)

Summary

User reports Gemma 4 26B at 75 t/s generation and 1500 t/s prompt processing on 2x RTX 4060 8GB. Setup is LM Studio with an IQ quant, serving to Hermes. User credits mradermacher's IQ quants for getting the model running.