- quant:
- 4bit
User runs Qwen3.6-35B-A3B-4bit on an M3 Max 128GB for production sub-agent delegations. User also mentions GLM-5.1 for orchestration. User is considering building a 5090 rig.
APPLE · 128GB unified memory · 2 reports
Use the calculator to check which models fit in 128 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.
How to compare these reports →All M3 Macs compared →User runs Qwen3.6-35B-A3B-4bit on an M3 Max 128GB for production sub-agent delegations. User also mentions GLM-5.1 for orchestration. User is considering building a 5090 rig.
M3 Max 128GB · MLX · 290,000 ctx
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 160 t/s prefill and 5-6 t/s generation on an M5 Max 128GB with Qwen 3.6 27B Q8 MLX at 290k context. GPU utilization sits at only 36-50%. User expected 8-14 t/s generation and asks how other setups compare.