Subscribe
16:57Anthropic names seven Chinese labs, 189.9m distilled exchanges and a Russia-linked actor.15:42Revised CLARITY Act runs 630 pages; stablecoin yield section unchanged from July14:35OpenAI launches ChatGPT for Financial Services with Daloopa and PitchBook data built in.14:17OpenAI pauses new $200 ChatGPT Pro sign-ups, seven days after GPT-6 Astra shipped.14:00Coinbase renames Base App back to Coinbase Wallet, adds Robinhood Chain and Monad13:54AMD launches Ryzen 5 5500F at $99 and Ryzen 5 7500 at $189 as DDR4 spot inverts

DeepSeek's V4.1-Flash costs $0.15 per million tokens for four-fifths of the week

The MIT-licensed release ties Opus 5 on DeepSWE and trails it by 20.6 points on Terminal-Bench 4.0, and the checkpoint is 38% bigger than the number in the headlines.

In briefDeepSeek announced V4.1-Flash on September 10 as the smallest model in a new architecture family with native visual understanding1Peak pricing is $0.30 input and $1.20 output per million tokens; off-peak is half; cached input is $0.006 peak and $0.003 off-peak2Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays; all other hours are off-peak at half price3
The Hangzhou skyline across West Lake, DeepSeek’s home city
Photo: Y Chen (CC BY-SA 4.0)

DeepSeek released V4.1-Flash on September 10 at $0.30 per million input tokens and $1.20 per million output, MIT-licensed, with a million-token context and native vision. Then read the footnote on the price page. Peak hours run 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, and everything outside them costs half. That is 35 hours out of 168, so the advertised price applies to 20.8% of the week and the real rate for most of the world most of the time is $0.15 in and $0.60 out. Cached input off-peak is $0.003 per million.

Which of those is the real price? The off-peak one, for anybody outside Chinese office hours — which is to say most of the people reading this. It is less a price than a rounding error with a decimal point in it.

Artificial Analysis put the model at 40 on its Intelligence Index, ahead of DeepSeek's own V4 Pro, a 1.6-trillion-parameter model, at roughly a quarter of the cost per token. Keep their cost-per-task figure.

Cost per Artificial Analysis Intelligence Index task ($)
GLM-5.32.01Kimi K32DeepSeek V4 Pro0.67DeepSeek V4.1 Flash0.27

We spent the afternoon in the model card and the repository rather than the launch thread, and two things there are worth your time.

Start with the parameter count. Everyone is repeating 552B. But 552B is only the backbone. The card also lists Engram conditional memory at 196B parameters, "sparsely accessed via token-based lookup" — a lookup table the model reads instead of computing. Add the vision tower and the rest and the safetensors index on Hugging Face totals 763,205,315,794 parameters across 510 GB of weights. So the thing you download is about 38% larger than the number in every headline, and roughly a quarter of it is memory rather than computation. And that is the trick, a good one: the activated path is 8B parameters while the model reads your prompt and 16B while it writes, which is why input-heavy agent loops get cheap.

Then the benchmark table, which is more honest than its fans.

Terminal-Bench 4.0, pass@1 (DeepSeek's own comparison table)
Claude Opus 551.8GPT-5.6 Sol39.9GLM-5.337.9DeepSeek V4.1 Flash31.2Kimi K312.6

DeepSeek publishes that chart itself. On DeepSWE the model resolves 74.2% against Opus 5's 74.0%, which is a tie, and on Terminal-Bench 2.1 it leads the field at 90.6. But on Terminal-Bench 4.0 — the harder, newer one — Opus 5 scores 51.8 and V4.1-Flash 31.2, a gap of 20.6 points. On Humanity's Last Exam without tools it is 36.8 against Opus 5's 56.3. The agentic wins are real; the reasoning gap is real too, and only one of them is in the posts going around.

And there is a subtler cost. Artificial Analysis measured V4.1-Flash at 89,000 tokens per Intelligence Index task, the most verbose model they have on record, 62% more than DeepSeek's own V4 Pro and more than Fable 5.1 or Opus 5. A model that costs a quarter as much per token and talks three times as long is not a quarter as cheap. Even so, it comes out at $0.27 a task against $2.01 for GLM-5.3, so the arithmetic survives — but if you are budgeting agent runs, budget the tokens, not the rate card.

Our read: the price card is the release. Its architecture is a KV-cache paper (890 bytes per token, roughly a quarter of V4-Flash's) and the benchmark table is a strong second tier, and neither would matter much on its own. What matters is that a model tying Opus 5 on DeepSWE now costs $0.60 per million output tokens for five days and thirteen hours out of every seven. Somebody selling a coding agent at a flat monthly price has to decide whether to keep paying frontier rates for a difference their users cannot see on most tickets. We would expect at least two established coding products to add a DeepSeek route by the end of October, and we would take the other side of any claim that V4.1-Flash displaces Opus on hard, long-horizon work this year — the Terminal-Bench 4.0 gap is too wide.

Independent scoring lands in the same place. OpenDesign ran thirteen models through the same design tasks and gave Astra 82.7 and V4.1-Flash 81.2, at $1.61 and $0.023 per finished design respectively; eleven of the thirteen scored lower than DeepSeek and cost more.

One caveat we would like closed. Every code-agent number in the comparison table comes from DeepSeek's own harness in "Minimal mode" with a 1M-token window, which is a reasonable choice and also a choice. Somebody should rerun Terminal-Bench 4.0 on a neutral scaffold. Until then, the cheapest honest summary of this release is the one on the pricing page, not the one in the benchmark table.

Sources

01
DeepSeek announced V4.1-Flash on September 10 as the smallest model in a new architecture family with native visual understanding🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher…” — x.com · primary · Sep 10
02
Peak pricing is $0.30 input and $1.20 output per million tokens; off-peak is half; cached input is $0.006 peak and $0.003 off-peakPRICING (3) 1M INPUT TOKENS (CACHE HIT) OFF-PEAK $0.003 … PEAK $0.006 … 1M INPUT TOKENS (CACHE MISS) OFF-PEAK $0.15 … PEAK $0.3 … 1M OUTPUT TOKENS OFF-PEAK $0.6 … PEAK $1.2” — api-docs.deepseek.com · primary · Sep 10
03
Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays; all other hours are off-peak at half priceOff-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).” — api-docs.deepseek.com · primary · Sep 10
Show all 16 sources
04
Context length 1M, max output 384K, vision supportedCONTEXT LENGTH 1M MAX OUTPUT MAXIMUM: 384K … Vision ✓ Not supported” — api-docs.deepseek.com · primary · Sep 10
05
The model card describes a 552B-backbone MoE with Engram conditional memory of 196B parameters and 8B/16B activated parametersWe introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. […] This allows the model to activate only 8B parameters per token during…” — huggingface.co · primary · Sep 10
06
The Hugging Face safetensors index totals 763,205,315,794 parameters across roughly 510 GB of weights"safetensors":{"parameters":{"BF16":1976441856,"F32":42307282,"F8_E4M3":204015223296,"I8":557171343360},"total":763205315794,"sharded":true,"totalFileSize":510304178606}” — huggingface.co · primary · Sep 10
07
Global KV cache footprint is 890 bytes per token, roughly a quarter of DeepSeek-V4-FlashCombined with FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels), these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash.” — huggingface.co · primary · Sep 10
08
DeepSeek's own comparison table: Terminal-Bench 4.0 gives Opus 5 51.8 and V4.1-Flash 31.2; DeepSWE 74.0 vs 74.2; HLE 56.3 vs 36.8; Terminal-Bench 2.1 89.1 vs 90.6Terminal-Bench 4.0 (Pass@1) 51.8 39.9 12.6 37.9 12.4 7.0 31.2 … DeepSWE v1.1 (Resolved) 74.0 73.0 67.5 66.9 62.7 54.4 74.2 … HLE (Pass@1) 56.3 44.5 43.5 42.0† 42.7† 37.8† 36.8 … Terminal-Bench 2.1 (Pass@1) 89.1 88.8 88.3 88.2 87.9 82.7 90.6” — huggingface.co · primary · Sep 10
09
Code-agent benchmarks were run in DeepSeek's own harness in Minimal mode with a 1M-token contextFor code agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench), the model is evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window.” — huggingface.co · primary · Sep 10
10
Artificial Analysis scored V4.1-Flash at 40 on its Intelligence Index, ahead of the 1.6T V4 Pro at ~4x less per tokenDeepSeek V4.1 Flash overtakes DeepSeek V4 Pro 0813 as DeepSeek’s new flagship model with a score of 40 on Artificial Analysis Intelligence Index. At just 552B parameters, it outperforms the Pro (1.6T) model while costing ~4x less per…” — x.com · primary · Sep 10
11
Artificial Analysis measured 89k tokens per Intelligence Index task, the most verbose they have recorded, at $0.27 per task against $2.01 for GLM-5.3, $2.00 for Kimi K3 and $0.67 for V4 ProDeepSeek V4.1 Flash is one of the most verbose models we’ve measured at 89k Tokens per Intelligence Index Task. That is 25% more than Z AI's flagship GLM-5.3 (71k) … Despite such verbosity, DeepSeek V4.1 Flash still costs just $0.27 per…” — x.com · primary · Sep 10
12
SemiAnalysis described the 552B backbone, 196B Engram memory and ~8x lower persistent KVCongrats to @deepseek_ai on releasing DeepSeek-V4.1-Flash! > 552B backbone, with a causal encoder-decoder activating just 8B params at prefill, 16B at decode > 196B Engram memory accessed through sparse lookups > ~8x less persistent KV…” — x.com · reported · Sep 10
13
vLLM served the model from day zero on NVIDIA and AMD GPUs and described Engram as 197B parameters of n-gram memory🐳 DeepSeek-V4.1-Flash is out, and vLLM serves it from day 0, verified on NVIDIA and AMD GPUs! 🎉 … ✨ Engram: a quarter of the checkpoint is n-gram memory the model looks up instead of computes. 197B parameters of it.” — x.com · reported · Sep 10
14
OpenDesign Arena scored V4.1-Flash 81.2 against GPT-6 Astra's 82.7, at $0.023 versus $1.61 per finished design, with 11 of 13 models scoring lower and costing moreOpenDesign Arena scored DeepSeek V4.1 Flash at 81.2 out of 100 on real-world design tasks, 98% of GPT-6 Astra's 82.7, while charging $0.023 per finished design against Astra's $1.61. Of the 13 models tested—including Claude Fable 5.1,…” — decrypt.co · reported · Sep 10
15
VentureBeat led on the $0.003/1M off-peak cached-input rateDeepSeek-V4.1-Flash debuts with $0.003/1M off-peak cached-input rate and benchmarks eclipsing GPT-5.6 Sol, Claude Opus 5” — x.com · reported · Sep 10
16
The model is MIT licensed and served from the first-party API➤ License: MIT ➤ Providers: DeepSeek first-party API” — x.com · primary · Sep 10
Up next · Keep readingAI · 4 min read

Anthropic's threat report: 189.9m distilled exchanges, and two Chinese labs that relayed their own users to Claude

The September 10 report names seven China-based labs, a Russian espionage actor whose agents rebuilt malware until it went undetected, and customer data that reached Anthropic by accident.

Continue ↓