Muse 30B Glimmer
M4 Max 128GB · WebGPU
- reported speed:
- 25.0 tokens/s generation
- quant:
- UD-Q2_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports running a model in-browser with custom WebGPU kernels, at speed comparable to llama.cpp.