Qwen3.6 35B (3B active)
on NVIDIA RTX 5090 Laptop 24GB · llama.cpp
Sep 20, 2026
Summary
User reports a DIY Jev-like inference setup using unmodified open weight LLMs, evaluated on a 32,235-example benchmark.
Qwen3.6 35B-A3B reached 75.5% accuracy at ~5.3 req/s on a laptop RTX 5090 24GB, with Qwen3 27B at 75.3% ~2.9 req/s and Qwen3-4B at 65.0% ~27 req/s.
The approach uses boolean verification of candidate answers via true/false logits, batched through llama.cpp, with no NLI fine-tuning or classifier head. User notes the benchmark is not perfectly apples-to-apples and expresses skepticism about Jev hype.