Qwen3.5 4B
on M2 Max 64GB · Transformers
Sep 23, 2026
Summary
User reports Qwen3.5 4B at 70.4 t/s on an M2 Max using native GGUF support in transformers.
Setup is the Q4_K_M GGUF from unsloth, reusing ggml kernels on Apple Silicon. llama.cpp reached 71.8 t/s on the same model.
Two other models were also tested: Qwen3.8 27B UD-Q4_K_M at 15.9 t/s (vs 13.4 t/s with llama.cpp) and Qwen3.5 35B-A3B UD-IQ4_XS at 60.2 t/s (vs 61.3 t/s).