Agents-A1 35B (3B active) Millie
iPhone 17 Pro
- reported speed:
- 22.0 tokens/s generation · 261.0 tokens/s prompt processing
- quant:
- ternary
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Millie is a series of compressed agentic models derived from Agents-A1 (a Qwen 3.5 35B-A3B finetune). The ternary expert model got 56% on SWE-bench Verified and runs at 22 tokens/s decode and 261 tokens/s pre-fill on an iPhone 17 Pro (12 GB RAM). The 2-bit expert version got 60%. The software targets Macs with 16 GB+ memory and Linux gaming PCs with 16 GB+ system RAM and as little as 4 GB VRAM. The harness is forked from OpenAI Codex. The user is seeking testers for AMD GPUs and small 4-8 GB cards.