- reported speed:
- 40.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports running 3 local sessions of GLM 4.7 at about 40 tok/s each on an AMD Strix Halo mini PC with 128GB unified RAM. The model family is inferred as GLM-4.7 from the post title and text. The user is positive about the setup for budget local AI.