Mellum2.1 12B (2.5B active)
on Unknown GPU · llama.cpp · 131,072 ctx
Oct 11, 2026
Use cases
codingagentictool-use
Summary
User reports Mellum2.1-12B-A2.5B at around 40 t/s on a laptop with 8GB VRAM and 32GB RAM.
Setup is llama.cpp (llama-server) with Q8 at 131K context, run through the Pi coding agent.
User says agentic behavior and tool calling are good, and the model often catches and fixes its own editing mistakes, but one-shot project results were mixed: Pelican SVG, Browser OS and Minecraft were poor, Bouncing Hexagon had okay physics but ran in the terminal, and Flappy Bird was completed with basic visuals and too-difficult gameplay. It also explored an existing game project and changed the dogs' jump height to 3x, and converted screen recordings into GIFs. User notes it struggles with unfamiliar workflows and that Qwen3.6 35B-A3B gets slower with larger contexts and overheats the laptop.