llamaperf
Sep 20, 2026
VRAM reported
24 GB

Summary

User reports a DIY Jev-like inference setup using unmodified open weight LLMs, evaluated on a 32,235-example benchmark. Qwen3.6 35B-A3B reached 75.5% accuracy at ~5.3 req/s on a laptop RTX 5090 24GB, with Qwen3 27B at 75.3% ~2.9 req/s and Qwen3-4B at 65.0% ~27 req/s. The approach uses boolean verification of candidate answers via true/false logits, batched through llama.cpp, with no NLI fine-tuning or classifier head. User notes the benchmark is not perfectly apples-to-apples and expresses skepticism about Jev hype.