llamaperf

Phi-3.5

Microsoft · 1 report

As of 7 Oct 2026, Phi-3.5 3.8B at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

Phi-3.5 VRAM requirements by size and quant →

How does Phi-3.5 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Phi-3.5 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Phi-3.5 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Phi-3.5

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

Phi-3.5 3.8B mini

M4 Pro 64GB · mlx-lm · 4,096 ctx

Tone: positive
reported speed:
72.2 tokens/s generation
quant:
4bit (MLX)
kv:
kv4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Phi-3.5-mini-instruct-4bit at 72.2 t/s on a Mac Mini M4 Pro 64GB. Setup is mlx-lm 0.30.7 with 4-bit weights and kv4 KV-cache quantization at 4096 generated tokens. KV-cache quantization with kv4 uses 0.51 GB and is 1.1% faster than unquantized (71.4 t/s), while kv8 uses 0.91 GB and is 4.9% slower (67.9 t/s). User also tested Qwen2.5-14B-Instruct-4bit and found GQA halves KV-cache per token versus Phi-3.5-mini's MHA.

Sep 26, 2026

Get a weekly email of new Phi-3.5 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Phi-3.5 3.8B mini
M4 Pro 64GB
4bit
mlx-lm
4,09672.2 tokens/s