llamaperf

DeepSeek V4.1 Flash

2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.

Post is a defense of Artificial Analysis benchmarks, not a local-LLM hardware benchmark. No GPU, engine, quant, or t/s figures. Models mentioned: DeepSeek V4.1-Flash (552B, AA score 40), Qwen 3.8-Flash-Next (180B, AA score 40), GPT-6 Astra (Max), Fable 5.1. DeepSeek V4.1-Flash reportedly matches/exceeds Qwen 3.8-Flash-Next on most evaluations and beats GPT-6 Astra (Max) on AutomationBench-AA (agentic SaaS workflows), but falls behind on AA-Omniscience Non-Hallucination Rate. AA spent $13,129 to independently test Fable 5.1.

reported speed:
6.0 tokens/s generation · 30.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

CPU-only benchmark on a Xeon. ~30 TPS prompt processing and ~6 TPS generation. N-gram table offloaded on 50% of threads. Sloppily vibecoded by Opus 5.0, unreviewed. Goal is overnight/over-week agentic jobs.