llamaperf

Agents-A1

InternScience · 2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.
reported speed:
22.0 tokens/s generation · 261.0 tokens/s prompt processing
quant:
ternary

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

Millie is a series of compressed agentic models derived from Agents-A1 (a Qwen 3.5 35B-A3B finetune). The ternary expert model got 56% on SWE-bench Verified and runs at 22 tokens/s decode and 261 tokens/s pre-fill on an iPhone 17 Pro (12 GB RAM). The 2-bit expert version got 60%. The software targets Macs with 16 GB+ memory and Linux gaming PCs with 16 GB+ system RAM and as little as 4 GB VRAM. The harness is forked from OpenAI Codex. The user is seeking testers for AMD GPUs and small 4-8 GB cards.

Tone: positive
reported speed:
97.0 tokens/s generation
quant:
APEX Compact (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Model is a 35B MoE agentic model, beats Qwen3.6 and DeepSeek V4 Pro in some scenarios. Repository: SC117/Agents-A1-Uncensored-MTP-APEX-GGUF. Also mentions a 4B dense variant but not attractive.