Gemma 4 26B
2× NVIDIA RTX 4060 · LM Studio
- generation:
- 75.0 tokens/s
- prompt processing (prefill):
- 1500.0 tokens/s
- quant:
- IQ (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.
User reports Gemma 4 26B at 75 t/s generation and 1500 t/s prompt processing on 2x RTX 4060 8GB. Setup is LM Studio with an IQ quant, serving to Hermes. User credits mradermacher's IQ quants for getting the model running.