Mac setups · · 7 min read
M5 Ultra Mac Studio for local LLMs: early measurements
What owners measured on the new Mac Studio with M5 Max and M5 Ultra, how much faster the Ultra really is, and which memory size to buy for local models.
Practical ways to choose a local LLM, measure its speed and improve your setup.
Mac setups · · 7 min read
What owners measured on the new Mac Studio with M5 Max and M5 Ultra, how much faster the Ultra really is, and which memory size to buy for local models.
Hardware planning · · 7 min read
NVIDIA's RTX Spark brings 128 GB of unified memory to laptops. DGX Spark's measured speeds show what it will run and what NVIDIA hasn't said yet.
Measurement · · 5 min read
Get a clean output speed from Ollama's own response fields, and keep model loading and prompt reading out of the number you publish.
Reading benchmarks · · 6 min read
Two people with the same card can report wildly different tokens per second. Here's what to line up before you believe either number.
Choosing a model · · 6 min read
Pick a local model by whether it fits your memory and gets your real tasks right, then check its speed against reports from machines like yours.
Choosing a model · · 5 min read
Which engine should run your local model? Match it to how you'll use it, the formats you need and your hardware, then test two on the same prompt.
Measurement · · 5 min read
Log your GPU's power draw with nvidia-smi, line it up with a timed run and work out the energy each answer costs. Here's the method and the traps to avoid.
Hardware planning · · 5 min read
Serving a local model to your team or family? Here's how parallel slots work, why each one costs memory, and how to read per-user speed against the total.
Reading benchmarks · · 5 min read
A draft model or built-in MTP can make local answers arrive faster. Here's when it works, how to turn it on, and why your speed report should say it was on.
Hardware planning · · 6 min read
No graphics card? You can still run a local model. Here's why memory bandwidth sets your speed, and how to test a CPU-only setup before you buy anything.
Measurement · · 6 min read
Time to first token is the wait before any text appears. Measure it from a streaming request and keep model loading and prompt caching out of it.
Choosing a model · · 6 min read
A mixture-of-experts model uses only a few experts per token, yet all of them have to sit in memory. Here's what that means for your hardware.
Troubleshooting · · 5 min read
What happens when a model spills past your GPU memory, how to set llama.cpp's -ngl, how to read Ollama's CPU/GPU split and what to test.
Hardware planning · · 5 min read
Every token in your context keeps its keys and values in memory. Here's how that cache grows, how to size it and how to shrink it in Ollama.
Quantization · · 6 min read
Q4_K_M, Q6_K, Q8_0, IQ4_XS: what each part of a GGUF file name tells you about bits per weight, file size and the memory you'll need.