llamaperf
Sep 18, 2026
Throughput
50.0 t/s gen
Quant
Q8

Use cases

coding

Summary

User reports Qwen3.8 27B at around 50 t/s on a CMP 170HX, running at Q8 or BF16. The user finds the model hallucinates classes and APIs in 90%+ of answers to domain-specific enterprise Java questions, while DeepSeek and Microsoft Copilot produce working code 99% of the time. The user asks whether Qwen's agentic optimization is responsible and requests tips for a harness, system prompt, or skills to reduce hallucination.