DeepSeek V4.1 Flash launch: benchmarks vs Sol and Opus 5
DeepSeek launched V4.1 Flash on 10 September 2026 as a new 552B MoE, not a tune of V4-Flash. On DeepSeek’s own scoreboard it leads GPT-5.6 Sol and Claude Opus 5 on Terminal-Bench 2.1, DeepSWE, CyberGym, AutomationBench, and Agents’ Last Exam—while those flagships still win GPQA and the newer Terminal-Bench rows. AshnaAI is the place to pin that DeepSeek row next to Sol, Opus 5, GLM-5.3-Flash, and Ashna-X1.
AshnaAI

What DeepSeek shipped on 10 September
On 10 September 2026 DeepSeek published Introducing DeepSeek-V4.1-Flash. The company calls it the smallest model in a new architecture family, with native visual understanding, higher throughput, and benchmark results ahead of its own V4-Pro. Weights and the instruct scoreboard are on Hugging Face under an MIT license.
The first-party API id is `deepseek-flash`. DeepSeek retired V4-Flash and V4-Flash-Vision-Exp; `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` temporarily route to V4.1 Flash. From 04:00 UTC on 14 September 2026, `deepseek-v4-pro` requests also route to V4.1 Flash at Flash rates until a V4.1 Pro launches. That is a supply change for anyone still calling Pro by name.
A new base model, not a point release
The `.1` in the name looks like a tune. The model card is a rebuild. DeepSeek-V4.1-Flash is a 552-billion-parameter MoE trained from scratch on a 45T-token multimodal corpus. Prefill activates 8B parameters; decode activates 16B. V4-Flash was 284B / 13B. V4-Pro was 1.6T / 49B. The new Causal Encoder–Decoder is why a Flash SKU can post Pro-or-better agent scores while cutting the persistent KV cache to about one-quarter of V4-Flash’s HBM and one-eighth of its SSD.
That is the product claim for long-running agents: cache-hit charges often dominate the bill. Compressing KV is how DeepSeek undercuts flagship spend on input-heavy traces. It is also why a ChatGPT-only or Claude-only seat is a weak answer. Agent work still needs files, connectors, and a catalog the buyer can switch when the next lab ships a cheaper Flash that wins the terminal cells.
Official launch benchmarks vs Sol, Opus 5, and GLM
Read the table as DeepSeek’s launch scoreboard, not as a third-party bake-off. V4.1 Flash jumps Sol and Opus 5 on Terminal-Bench 2.1 (90.6 vs 88.8 / 89.1), DeepSWE v1.1 (74.2 vs 73.0 / 74.0), CyberGym (88.1 vs 84.5), AutomationBench (54.8 vs 45.8 / 50.3), and Agents’ Last Exam (31.8 vs 26.7 / 28.6). GPT-5.6 Sol still wins GPQA Diamond (94.1 vs 90.9). Claude Opus 5 still wins Terminal-Bench 4.0 (51.8 vs 31.2), ProgramBench (37.0 vs 20.3), NL2Repo-Bench (75.3 vs 64.0), and Humanity’s Last Exam without tools (56.3 vs 36.8). The newest DeepSeek badge is not a reason to retire every other pin.
The table below is the competitive instruct scoreboard from DeepSeek’s 10 September 2026 model card. V4.1 Flash leads several agent and terminal cells. It does not lead GPQA Diamond, Humanity’s Last Exam without tools, or the newer Terminal-Bench 3.0 and 4.0 rows.
| Benchmark | V4.1 Flash | GPT-5.6 Sol | Claude Opus 5 | GLM-5.3 | V4-Pro |
|---|---|---|---|---|---|
| Where V4.1 Flash leads | |||||
| Terminal-Bench 2.1 (Pass@1) | 90.6 | 88.8 | 89.1 | 88.2 | 87.9 |
| DeepSWE v1.1 (Resolved) | 74.2 | 73.0 | 74.0 | 66.9 | 62.7 |
| CyberGym (Pass@1) | 88.1 | 84.5 | — | 84.5 | 83.3 |
| AutomationBench (Pass@1) | 54.8 | 45.8 | 50.3 | 48.8 | 43.2 |
| Agents' Last Exam (Pass@1) | 31.8 | 26.7 | 28.6 | 28.5 | 25.7 |
| HLE with tools (Pass@1) | 63.9 | — | 63.6 | 62.5 | 60.0 |
| Codeforces (rating) | 3471 | — | — | — | 3348 |
| Where flagships still lead | |||||
| GPQA Diamond (Pass@1) | 90.9 | 94.1 | 93.4 | 88.1 | 92.4 |
| Humanity's Last Exam (Pass@1) | 36.8 (39.1*) | 44.5 | 56.3 | 42.0* | 42.7* |
| Terminal-Bench 3.0 (Pass@1) | 30.0 | 34.4 | 43.3 | 28.3 | 11.8 |
| Terminal-Bench 4.0 (Pass@1) | 31.2 | 39.9 | 51.8 | 37.9 | 12.4 |
| ProgramBench (Almost@1) | 20.3 | 23.0 | 37.0 | 19.0 | 15.5 |
| NL2Repo-Bench (score) | 64.0 | 56.8 | 75.3 | 58.0 | 61.5 |
| Vision with tools | |||||
| Chartography with tools (Pass@1) | 78.9 | 79.9 | 84.0 | — | — |
| BabyVision with tools (Pass@1) | 89.6 | 88.9 | 94.1 | — | — |
| ZeroBench-main with tools (Pass@5) | 49.0 | 53.0 | 52.0 | — | — |
What the token bill looks like
V4.1 Flash off-peak is $0.15 cache-miss input and $0.60 output per 1M tokens. Peak doubles that. On the AshnaAI list rates already used for GLM-5.3-Flash vs GPT-5.6 Sol and GLM-5.3-Flash vs Claude Opus 5, Sol is about $5 / $30 and Opus 5 is $5 / $25. Cache-miss input matches GLM-5.3-Flash; output is a few cents higher. Volume coding that does not need V4.1’s agent jump should stay on Flash. Mixed everyday work stays on Ashna-X1.
| Rate per 1M tokens | V4.1 Flash (off-peak) | V4.1 Flash (peak) | GPT-5.6 Sol | Claude Opus 5 | GLM-5.3-Flash |
|---|---|---|---|---|---|
| Input (cache hit) | $0.003 | $0.006 | — | — | — |
| Input (cache miss) | $0.15 | $0.30 | ~$5.00 | $5.00 | $0.15 |
| Output | $0.60 | $1.20 | ~$30.00 | $25.00 | $0.50 |
How to read the wins and the gaps
The shape is coherent. V4.1 Flash is strongest on the coding, terminal, and automation work the new architecture and agent post-training were built for. It is weaker on specialist-knowledge and the newest Terminal-Bench generations, where Opus 5 still has a 20-point gap on 4.0 and a 19-point gap on Humanity’s Last Exam without tools. DeepSeek reports HLE with tools at 63.9, a hair above Opus 5’s 63.6—so the tool-using agent story is closer than the no-tools science story.
Harness variance is real. DeepSeek’s own scaffold table shows DeepSWE v1.1 moving from 65.5 on OpenCode to 74.2 on mini-SWE at the same max effort. Treat every cell as approximate. A 0.2-point DeepSWE lead over Opus 5 is not a reason to delete the Anthropic pin. A 54.8 vs 45.8 AutomationBench lead, at Flash-tier list, is a reason to stop sending every agent trace to Sol.
How AshnaAI is the better place to run this launch
DeepSeek’s app gives you DeepSeek’s product surface. AshnaAI gives you the catalog: pin DeepSeek V4 Flash today, pin V4.1 Flash when that row is live, keep GPT-5.6 Sol and Claude Opus 5 for the GPQA and Terminal-Bench 4.0 cells they still win, and leave GLM-5.3-Flash as the coding default so DeepSeek tokens stay reserved for long agent traces.
The same ids run through the OpenAI-compatible API. Create a key in Account → API, then follow How to call any catalog model through the API. Cost-quality ladder: Cheapest coding AI vs Claude Opus 5 and GPT-5.6 Sol. Routing that already starts simple asks on DeepSeek V4 Flash: Reduce LLM API cost by 95%. Compare buying criteria: AshnaAI vs ChatGPT and AshnaAI vs Claude.
Try the catalog now
New accounts start at app.ashna.ai/signup. Open AshnaAI chat and pin DeepSeek V4 Flash for the current DeepSeek agent row, or GLM-5.3-Flash for volume coding. When the catalog shows DeepSeek V4.1 Flash, pin it the same way you pin Flash. Do not wait for a DeepSeek-only seat to decide which model your team can use.
Product API traffic uses Account → API and the AshnaAI API docs.
Frequently asked questions
- What is DeepSeek V4.1 Flash?
- DeepSeek-V4.1-Flash is DeepSeek’s 10 September 2026 production model. It is a 552-billion-parameter multimodal mixture-of-experts with a new Causal Encoder–Decoder architecture, native image understanding, a 1-million-token context, and MIT-licensed weights. DeepSeek’s API name is deepseek-flash.
- Does DeepSeek V4.1 Flash beat GPT-5.6 Sol and Claude Opus 5?
- On DeepSeek’s published max-effort table, yes on several agent and coding cells: Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, AutomationBench, Agents’ Last Exam, and HLE with tools. No on the hardest knowledge and newer terminal rows: GPQA Diamond, Humanity’s Last Exam without tools, Terminal-Bench 3.0 and 4.0, and ProgramBench. Those are vendor-published numbers, not an independent bake-off.
- Is V4.1 Flash just a faster V4-Flash?
- No. DeepSeek filed it as the smallest member of a new architecture family: 552B backbone versus 284B for V4-Flash, 8B/16B active versus 13B, and a causal encoder–decoder instead of the prior stack. DeepSeek also says third-party tests put V4.1 Flash ahead of its own V4-Pro on performance, cost, speed, and total runtime, and it will route deepseek-v4-pro traffic to V4.1 Flash from 04:00 UTC on 14 September 2026 until a V4.1 Pro exists.
- How much does DeepSeek V4.1 Flash cost?
- DeepSeek’s launch pricing (USD per 1M tokens) is $0.003 cache-hit input, $0.15 cache-miss input, and $0.60 output off-peak. Peak rates are 2× that. Cache-miss input matches GLM-5.3-Flash’s $0.15 list; output is $0.60 versus Flash’s $0.50. GPT-5.6 Sol remains about $5 / $30 and Claude Opus 5 is $5 / $25.
- What is new in the architecture?
- A 40-layer Causal Encoder–Decoder (20 encoder layers, 20 decoder layers) so the decoder’s global KV cache is projected from the encoder instead of stored per decoder layer. Combined with Compressed Sparse Attention 2 and FP4 KV caching, DeepSeek reports about 1/4 the HBM and 1/8 the SSD footprint versus V4-Flash. That is why long agent traces get cheaper, not only the list rate.
- Where do I try DeepSeek V4.1 Flash without locking into DeepSeek’s app?
- Open AshnaAI at https://app.ashna.ai/signup Pin DeepSeek V4 Flash today at https://app.ashna.ai/chat?agent=deepseek-v4-flash keep GLM-5.3-Flash for volume coding at https://app.ashna.ai/chat?agent=glm-5.3-flash and pin the V4.1 Flash row the same way when that catalog id is live. The same ids go through the OpenAI-compatible API.
Tags
Related
- AshnaAI for Work
- AshnaAI vs ChatGPT
- AshnaAI vs Claude
- Large Language Model (LLM)
- Foundation Model
- gpt 6 astra launch benchmarks
- glm 5 3 flash vs gpt 5 6 sol
- glm 5 3 flash vs claude opus 5
- glm 5 3 flash vs kimi k3
- pin glm 5 3 flash for coding
- how to call any catalog model through the api
- cheapest coding ai vs claude opus 5 and gpt 5 6 sol
- reduce llm api cost 95 percent
- how to use ashna x1 instead of picking models yourself
- ashna x1 task aware model routing
Try this in AshnaAI. Create a free account.
Found this article helpful? Share it with your network.