Maple Preview 20B (1B active)
M4 16GB · Mference · 131,072 ctx
- reported speed:
- 20.0 tokens/s generation · 40.0 tokens/s prompt processing
- quant:
- ternary
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Maple Preview, a 20B A1B MoE model trained from scratch in ternary precision, running on a MacBook Air M4 with 16 GB RAM. Setup is Mference, a fork of turbo-fieldfare, streaming experts from SSD and reducing memory usage to 500-1200 MB. The model has limited world knowledge and multilingual capabilities but can use tools. The user is enthusiastic about the low memory footprint and potential for background agentic workloads.