llamaperf

Qwen3 14B Heretic

on NVIDIA RTX 5090 Laptop 24GB · OpenVINO GenAI

Tone: mixed
Sep 18, 2026
Throughput
3.6 t/s gen
Quant
INT4_SYM (OpenVINO IR)
System RAM
32 GB
VRAM reported
24 GB

Summary

User reports Qwen3-14B Heretic at 3.56 t/s on an Intel AI Boost NPU37XX in an Acer Predator Helios 16 AI laptop with 32 GB RAM and an RTX 5090 Laptop 24GB. Setup is OpenVINO GenAI with the machine-made-Fibre INT4_SYM OpenVINO IR export, group_size=-1, all_layers=true, loaded directly via openvino_genai.LLMPipeline on NPU with no conversion. Load took 33.52 s and 96 output tokens took 26.98 s; peak RAM was about 15.6 GB with a 9.7 GB working set. User also benchmarked Qwen3-8B INT4 SYM on the same NPU at 6.72 t/s with 6.15 s load, and found Qwen3-30B-A3B MoE impractical due to host memory pressure during OpenVINO preparation.