Guide

DeepSeek V4.1 Flash launch: benchmarks vs Sol and Opus 5

DeepSeek launched V4.1 Flash on 10 September 2026 as a new 552B MoE, not a tune of V4-Flash. On DeepSeek’s own scoreboard it leads GPT-5.6 Sol and Claude Opus 5 on Terminal-Bench 2.1, DeepSWE, CyberGym, AutomationBench, and Agents’ Last Exam—while those flagships still win GPQA and the newer Terminal-Bench rows. AshnaAI is the place to pin that DeepSeek row next to Sol, Opus 5, GLM-5.3-Flash, and Ashna-X1.

AshnaAI

DeepSeek V4.1 Flash launch: benchmarks vs Sol and Opus 5

What DeepSeek shipped on 10 September

On 10 September 2026 DeepSeek published Introducing DeepSeek-V4.1-Flash. The company calls it the smallest model in a new architecture family, with native visual understanding, higher throughput, and benchmark results ahead of its own V4-Pro. Weights and the instruct scoreboard are on Hugging Face under an MIT license.

The first-party API id is `deepseek-flash`. DeepSeek retired V4-Flash and V4-Flash-Vision-Exp; `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` temporarily route to V4.1 Flash. From 04:00 UTC on 14 September 2026, `deepseek-v4-pro` requests also route to V4.1 Flash at Flash rates until a V4.1 Pro launches. That is a supply change for anyone still calling Pro by name.

A new base model, not a point release

The `.1` in the name looks like a tune. The model card is a rebuild. DeepSeek-V4.1-Flash is a 552-billion-parameter MoE trained from scratch on a 45T-token multimodal corpus. Prefill activates 8B parameters; decode activates 16B. V4-Flash was 284B / 13B. V4-Pro was 1.6T / 49B. The new Causal Encoder–Decoder is why a Flash SKU can post Pro-or-better agent scores while cutting the persistent KV cache to about one-quarter of V4-Flash’s HBM and one-eighth of its SSD.

That is the product claim for long-running agents: cache-hit charges often dominate the bill. Compressing KV is how DeepSeek undercuts flagship spend on input-heavy traces. It is also why a ChatGPT-only or Claude-only seat is a weak answer. Agent work still needs files, connectors, and a catalog the buyer can switch when the next lab ships a cheaper Flash that wins the terminal cells.

Official launch benchmarks vs Sol, Opus 5, and GLM

Read the table as DeepSeek’s launch scoreboard, not as a third-party bake-off. V4.1 Flash jumps Sol and Opus 5 on Terminal-Bench 2.1 (90.6 vs 88.8 / 89.1), DeepSWE v1.1 (74.2 vs 73.0 / 74.0), CyberGym (88.1 vs 84.5), AutomationBench (54.8 vs 45.8 / 50.3), and Agents’ Last Exam (31.8 vs 26.7 / 28.6). GPT-5.6 Sol still wins GPQA Diamond (94.1 vs 90.9). Claude Opus 5 still wins Terminal-Bench 4.0 (51.8 vs 31.2), ProgramBench (37.0 vs 20.3), NL2Repo-Bench (75.3 vs 64.0), and Humanity’s Last Exam without tools (56.3 vs 36.8). The newest DeepSeek badge is not a reason to retire every other pin.

The table below is the competitive instruct scoreboard from DeepSeek’s 10 September 2026 model card. V4.1 Flash leads several agent and terminal cells. It does not lead GPQA Diamond, Humanity’s Last Exam without tools, or the newer Terminal-Bench 3.0 and 4.0 rows.

DeepSeek-V4.1-Flash competitive benchmarks published on 10 September 2026
BenchmarkV4.1 FlashGPT-5.6 SolClaude Opus 5GLM-5.3V4-Pro
Where V4.1 Flash leads
Terminal-Bench 2.1 (Pass@1)90.688.889.188.287.9
DeepSWE v1.1 (Resolved)74.273.074.066.962.7
CyberGym (Pass@1)88.184.584.583.3
AutomationBench (Pass@1)54.845.850.348.843.2
Agents' Last Exam (Pass@1)31.826.728.628.525.7
HLE with tools (Pass@1)63.963.662.560.0
Codeforces (rating)34713348
Where flagships still lead
GPQA Diamond (Pass@1)90.994.193.488.192.4
Humanity's Last Exam (Pass@1)36.8 (39.1*)44.556.342.0*42.7*
Terminal-Bench 3.0 (Pass@1)30.034.443.328.311.8
Terminal-Bench 4.0 (Pass@1)31.239.951.837.912.4
ProgramBench (Almost@1)20.323.037.019.015.5
NL2Repo-Bench (score)64.056.875.358.061.5
Vision with tools
Chartography with tools (Pass@1)78.979.984.0
BabyVision with tools (Pass@1)89.688.994.1
ZeroBench-main with tools (Pass@5)49.053.052.0
Copied from DeepSeek-AI’s DeepSeek-V4.1-Flash model card instruct table. All rows use maximum reasoning effort. Agent coding benches use DeepSeek Harness Minimal mode and a 1M-token context. HLE 36.8 is the full set; 39.1 is the text-only subset. Blank cells were unpublished. huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

What the token bill looks like

V4.1 Flash off-peak is $0.15 cache-miss input and $0.60 output per 1M tokens. Peak doubles that. On the AshnaAI list rates already used for GLM-5.3-Flash vs GPT-5.6 Sol and GLM-5.3-Flash vs Claude Opus 5, Sol is about $5 / $30 and Opus 5 is $5 / $25. Cache-miss input matches GLM-5.3-Flash; output is a few cents higher. Volume coding that does not need V4.1’s agent jump should stay on Flash. Mixed everyday work stays on Ashna-X1.

List rates that decide whether DeepSeek V4.1 Flash is a daily pin
Rate per 1M tokensV4.1 Flash (off-peak)V4.1 Flash (peak)GPT-5.6 SolClaude Opus 5GLM-5.3-Flash
Input (cache hit)$0.003$0.006
Input (cache miss)$0.15$0.30~$5.00$5.00$0.15
Output$0.60$1.20~$30.00$25.00$0.50
V4.1 Flash peak and off-peak rates from DeepSeek’s 10 September 2026 launch note and API pricing page. Sol, Opus 5, and GLM-5.3-Flash are the AshnaAI catalog list rates used on the existing model-guide pages. www.deepseek.com/en/news/deepseek-v4-1-flash

How to read the wins and the gaps

The shape is coherent. V4.1 Flash is strongest on the coding, terminal, and automation work the new architecture and agent post-training were built for. It is weaker on specialist-knowledge and the newest Terminal-Bench generations, where Opus 5 still has a 20-point gap on 4.0 and a 19-point gap on Humanity’s Last Exam without tools. DeepSeek reports HLE with tools at 63.9, a hair above Opus 5’s 63.6—so the tool-using agent story is closer than the no-tools science story.

Harness variance is real. DeepSeek’s own scaffold table shows DeepSWE v1.1 moving from 65.5 on OpenCode to 74.2 on mini-SWE at the same max effort. Treat every cell as approximate. A 0.2-point DeepSWE lead over Opus 5 is not a reason to delete the Anthropic pin. A 54.8 vs 45.8 AutomationBench lead, at Flash-tier list, is a reason to stop sending every agent trace to Sol.

How AshnaAI is the better place to run this launch

DeepSeek’s app gives you DeepSeek’s product surface. AshnaAI gives you the catalog: pin DeepSeek V4 Flash today, pin V4.1 Flash when that row is live, keep GPT-5.6 Sol and Claude Opus 5 for the GPQA and Terminal-Bench 4.0 cells they still win, and leave GLM-5.3-Flash as the coding default so DeepSeek tokens stay reserved for long agent traces.

The same ids run through the OpenAI-compatible API. Create a key in Account → API, then follow How to call any catalog model through the API. Cost-quality ladder: Cheapest coding AI vs Claude Opus 5 and GPT-5.6 Sol. Routing that already starts simple asks on DeepSeek V4 Flash: Reduce LLM API cost by 95%. Compare buying criteria: AshnaAI vs ChatGPT and AshnaAI vs Claude.

Try the catalog now

New accounts start at app.ashna.ai/signup. Open AshnaAI chat and pin DeepSeek V4 Flash for the current DeepSeek agent row, or GLM-5.3-Flash for volume coding. When the catalog shows DeepSeek V4.1 Flash, pin it the same way you pin Flash. Do not wait for a DeepSeek-only seat to decide which model your team can use.

Product API traffic uses Account → API and the AshnaAI API docs.

Frequently asked questions

What is DeepSeek V4.1 Flash?
DeepSeek-V4.1-Flash is DeepSeek’s 10 September 2026 production model. It is a 552-billion-parameter multimodal mixture-of-experts with a new Causal Encoder–Decoder architecture, native image understanding, a 1-million-token context, and MIT-licensed weights. DeepSeek’s API name is deepseek-flash.
Does DeepSeek V4.1 Flash beat GPT-5.6 Sol and Claude Opus 5?
On DeepSeek’s published max-effort table, yes on several agent and coding cells: Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, AutomationBench, Agents’ Last Exam, and HLE with tools. No on the hardest knowledge and newer terminal rows: GPQA Diamond, Humanity’s Last Exam without tools, Terminal-Bench 3.0 and 4.0, and ProgramBench. Those are vendor-published numbers, not an independent bake-off.
Is V4.1 Flash just a faster V4-Flash?
No. DeepSeek filed it as the smallest member of a new architecture family: 552B backbone versus 284B for V4-Flash, 8B/16B active versus 13B, and a causal encoder–decoder instead of the prior stack. DeepSeek also says third-party tests put V4.1 Flash ahead of its own V4-Pro on performance, cost, speed, and total runtime, and it will route deepseek-v4-pro traffic to V4.1 Flash from 04:00 UTC on 14 September 2026 until a V4.1 Pro exists.
How much does DeepSeek V4.1 Flash cost?
DeepSeek’s launch pricing (USD per 1M tokens) is $0.003 cache-hit input, $0.15 cache-miss input, and $0.60 output off-peak. Peak rates are 2× that. Cache-miss input matches GLM-5.3-Flash’s $0.15 list; output is $0.60 versus Flash’s $0.50. GPT-5.6 Sol remains about $5 / $30 and Claude Opus 5 is $5 / $25.
What is new in the architecture?
A 40-layer Causal Encoder–Decoder (20 encoder layers, 20 decoder layers) so the decoder’s global KV cache is projected from the encoder instead of stored per decoder layer. Combined with Compressed Sparse Attention 2 and FP4 KV caching, DeepSeek reports about 1/4 the HBM and 1/8 the SSD footprint versus V4-Flash. That is why long agent traces get cheaper, not only the list rate.
Where do I try DeepSeek V4.1 Flash without locking into DeepSeek’s app?
Open AshnaAI at https://app.ashna.ai/signup Pin DeepSeek V4 Flash today at https://app.ashna.ai/chat?agent=deepseek-v4-flash keep GLM-5.3-Flash for volume coding at https://app.ashna.ai/chat?agent=glm-5.3-flash and pin the V4.1 Flash row the same way when that catalog id is live. The same ids go through the OpenAI-compatible API.

Tags

#DeepSeek V4.1 Flash#DeepSeek#benchmarks#multi-model

Found this article helpful? Share it with your network.