llamaperf

Qwen3.5 4B

on M2 Max 64GB · Transformers

Sep 23, 2026
Throughput
70.4 t/s gen
Quant
Q4_K_M (GGUF)
System RAM
64 GB

Summary

User reports Qwen3.5 4B at 70.4 t/s on an M2 Max using native GGUF support in transformers. Setup is the Q4_K_M GGUF from unsloth, reusing ggml kernels on Apple Silicon. llama.cpp reached 71.8 t/s on the same model. Two other models were also tested: Qwen3.8 27B UD-Q4_K_M at 15.9 t/s (vs 13.4 t/s with llama.cpp) and Qwen3.5 35B-A3B UD-IQ4_XS at 60.2 t/s (vs 61.3 t/s).