Mellum2.1 12B (2.5B active)
Unknown GPU · llama.cpp · 131,072 ctx
- generation:
- 40.0 tokens/s
- quant:
- Q8 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.
User reports Mellum2.1-12B-A2.5B at around 40 t/s on a laptop with 8GB VRAM and 32GB RAM. Setup is llama.cpp (llama-server) with Q8 at 131K context, run through the Pi coding agent. User says agentic behavior and tool calling are good, and the model often catches and fixes its own editing mistakes, but one-shot project results were mixed: Pelican SVG, Browser OS and Minecraft were poor, Bouncing Hexagon had okay physics but ran in the terminal, and Flappy Bird was completed with basic visuals and too-difficult gameplay. It also explored an existing game project and changed the dogs' jump height to 3x, and converted screen recordings into GIFs. User notes it struggles with unfamiliar workflows and that Qwen3.6 35B-A3B gets slower with larger contexts and overheats the laptop.