GLM-4.5-Air 106B (12B active)
RTX 3090 · llama.cpp
- mtp (multi-token prediction):
- on
creative-writingrp
User recommends enabling MTP in llama.cpp for GLM-4.5-Air, a 106B MoE with 12B active params. Runs on 3090s. Mentions creative writing/RP finetunes and Intellect 3.x. Also notes MTP works for full GLM-4.5.