DeepSeek Just Launched V4.1 Flash — And Says It Beats Its Own Pro Model
The Chinese AI challenger is once again attacking the industry’s economics by combining stronger performance with some of the lowest API prices available.
As OpenAI, Anthropic and Google race towards increasingly powerful systems, DeepSeek has returned with another aggressive challenge on capability and cost.
DeepSeek has officially launched DeepSeek-V4.1-Flash, a new multimodal artificial intelligence model built on an entirely new architecture that the Chinese company says is faster, cheaper and more capable than its existing V4 Pro model.
The release on September 10 represents considerably more than another routine model update. DeepSeek describes V4.1 Flash as the smallest member of a new architecture family designed for a higher capability ceiling, faster inference, greater throughput and eventual scaling into significantly larger models. The company has also built native visual understanding directly into the system, allowing it to process images as well as text.
Perhaps the most striking claim, however, is that DeepSeek believes the new Flash model has already surpassed DeepSeek V4 Pro across performance, cost, speed and total task-completion time.
That claim is strong enough that DeepSeek plans to redirect requests for its existing Pro model towards V4.1 Flash before the eventual arrival of V4.1 Pro.
For an AI industry increasingly obsessed not only with intelligence but with the cost of delivering it, that could prove important.
DeepSeek V4.1 Flash Is Built on a Completely New Architecture
DeepSeek says V4.1 Flash is the first publicly released member of a broader new architecture family.
The company says the design was built around four priorities: raising the model's potential capability ceiling, increasing inference speed, improving throughput and providing an architecture capable of scaling towards substantially larger future models.
That final point may ultimately matter more than V4.1 Flash itself.
Flash appears to be a demonstration of the foundations DeepSeek intends to use for future systems rather than merely an isolated product.
According to details shared alongside the release, V4.1 Flash is a 552-billion-parameter mixture-of-experts model using what DeepSeek describes as a Causal-Encoder-Decoder architecture.
But all 552 billion parameters are not running for every token.
DeepSeek says only around 8 billion parameters are activated on the input side and 16 billion on the output side. That asymmetric structure is designed to reduce the computational burden of operating such a large model while retaining access to much greater underlying capacity.
This is one of the most important battles currently taking place in artificial intelligence.
Building a highly capable model is only part of the challenge. Companies must also make those models fast enough and inexpensive enough that millions of people and businesses can actually use them.
DeepSeek built much of its international reputation around precisely that problem.
V4.1 Flash suggests it intends to keep pushing in the same direction.
DeepSeek Says Flash Has Already Beaten V4 Pro
The naming could initially make V4.1 Flash sound like a lightweight alternative to DeepSeek's more expensive Pro system.
DeepSeek's own assessment suggests almost the opposite.
The company says extensive internal and external testing found that V4.1 Flash surpassed V4 Pro across four important measures: performance, cost, speed and overall task-completion time.
DeepSeek has consequently decided that continuing to serve V4 Pro requests at higher costs while V4.1 Flash performs better would make little sense.
V4 Pro is therefore being phased out, with its shutdown now scheduled for September 14, according to reports citing DeepSeek API notifications. Requests will then be routed towards V4.1 Flash while users wait for the future V4.1 Pro.
That raises an obvious question.
If Flash is already this strong, what exactly is DeepSeek preparing for V4.1 Pro?
DeepSeek has not yet given the complete answer.
But the architecture's explicit ability to scale towards larger models provides one of the clearest clues.
V4.1 Flash Can Natively Understand Images
Another significant change is multimodality.
DeepSeek V4.1 Flash provides native multimodal visual understanding, meaning image comprehension is incorporated into the model rather than being treated purely as an external addition.
That expands the range of tasks the model can potentially perform.
A multimodal AI can inspect screenshots, interpret charts, analyse documents containing visual information, reason about photographs and combine what it sees with text instructions.
Those abilities become particularly important as AI shifts away from simple chatbot conversations towards autonomous agents capable of interacting with computers and digital environments.
Modern AI agents may need to look at an interface, determine what is happening on-screen, understand written instructions and then decide what action to perform.
Vision therefore becomes increasingly central to the broader AI-agent race.
Coding and Agent Benchmarks Show a Major Jump
DeepSeek has published benchmark results suggesting V4.1 Flash is particularly strong across reasoning, software engineering, cybersecurity and agentic tasks.
Among the figures released by the company are a 90.9 score on GPQA Diamond, a Codeforces rating of 3471, 90.6 on Terminal-Bench 2.1, 74.2 on DeepSWE v1.1, 88.1 on CyberGym and 54.8 on Automation-Bench.
Benchmarks should always be interpreted carefully.
Individual tests capture particular capabilities under particular conditions, and strong benchmark performance does not guarantee that a model will be superior for every real-world task.
But the direction of travel is significant.
DeepSeek's previous V4 Flash release scored 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE according to the company's published API changelog. V4.1 Flash's reported scores of 90.6 and 74.2 respectively therefore indicate a substantial improvement on those tests.
The gains are particularly relevant because coding agents are emerging as one of the most commercially important applications of generative AI.
These systems do not simply suggest individual lines of code. Increasingly capable models can inspect repositories, plan modifications, use development tools, run commands, diagnose failures and work through longer software-engineering tasks.
DeepSeek clearly wants V4.1 Flash competing in that market.
DeepSeek Is Cutting the Cost at the Same Time
Capability is only half of the story.
DeepSeek has simultaneously introduced new pricing for its Flash series.
Reported direct API pricing during off-peak periods is approximately $0.003 per million cached input tokens, $0.15 per million uncached input tokens and $0.60 per million output tokens.
Peak-hour prices are double those figures.
In Chinese yuan, the announced off-peak rates are RMB 0.02 for cached input, RMB 1 for uncached input and RMB 4 for output per million tokens.
That combination — stronger performance and falling inference costs — is potentially more disruptive than benchmark improvements alone.
Businesses deploying AI at scale can process enormous numbers of tokens.
Small differences in price therefore become substantial when multiplied across millions of users, customer-service interactions, coding tasks or autonomous-agent actions.
An inexpensive model that is merely adequate can already be commercially useful.
An inexpensive model that begins approaching or overtaking premium alternatives is considerably more threatening to competitors.
DeepSeek Is Once Again Attacking the Economics of AI
DeepSeek became one of the most closely watched companies in artificial intelligence because it challenged an assumption that had dominated much of the industry: that frontier-level capabilities necessarily required extraordinarily expensive models to operate.
V4.1 Flash continues that strategy.
Rather than competing solely by saying its model is smarter, DeepSeek is emphasising the relationship between intelligence, speed and price.
That is becoming increasingly important as AI transitions from occasional chatbot usage towards systems that may perform hundreds or thousands of actions on behalf of individual users.
An autonomous agent researching information, analysing files, writing software and repeatedly calling external tools can consume vastly more inference than a conventional chatbot answering one question.
The price of intelligence therefore becomes part of the capability itself.
A slightly weaker model that can economically perform 1,000 actions may sometimes be more useful than an exceptionally powerful model that is too expensive to run continuously.
DeepSeek appears determined to exploit that equation.
The Launch Comes During an Intensifying Global AI Race
The timing is also notable.
DeepSeek's release arrives during an extraordinary period of competition between Chinese and American AI laboratories.
OpenAI recently launched GPT-6 Astra, while Anthropic and Google have continued pushing their own increasingly capable systems. The race is expanding beyond raw chatbot intelligence into coding, autonomous agents, cybersecurity, multimodal reasoning and increasingly long-running tasks.
DeepSeek is also reportedly preparing for an initial public offering on Shanghai's STAR Market.
Reuters reported that the company has engaged CITIC Securities as it prepares for a possible domestic listing and seeks additional funding for computing infrastructure, model development and retaining highly sought-after AI researchers.
That makes V4.1 Flash commercially significant as well as technically interesting.
DeepSeek is no longer simply trying to demonstrate that a Chinese laboratory can produce competitive AI.
It is attempting to build a sustainable company capable of financing the escalating cost of frontier-model development.
But DeepSeek Is Also Caught in a Growing US-China Dispute
The release arrives against an increasingly contentious political backdrop.
Just two days before the V4.1 Flash launch, the United States accused six Chinese artificial-intelligence companies — including DeepSeek — of improperly using outputs from American AI systems to accelerate their own development through model distillation.
China rejected the accusations and argued that distillation is a widely used technological technique.
The dispute illustrates how strategically important advanced AI models have become.
The competition between laboratories is no longer purely commercial.
Governments increasingly view frontier artificial intelligence as infrastructure with potential implications for cybersecurity, defence, economic productivity and national power.
Every significant DeepSeek release therefore lands inside the wider technology rivalry between Washington and Beijing.
What Comes After V4.1 Flash Could Be Even More Important
V4.1 Flash itself is impressive primarily because of what DeepSeek says it can deliver: stronger capability, native vision, faster inference and unusually aggressive pricing.
But the deeper significance may be the architecture underneath it.
DeepSeek is explicitly describing Flash as the smallest model in a new family.
That wording strongly suggests larger versions are coming.
V4.1 Pro is already expected, and DeepSeek's architecture has been designed specifically to scale towards larger models.
If the smallest model in that family can genuinely outperform the company's previous Pro offering, the eventual larger systems become considerably more interesting.
The central question is therefore no longer simply whether V4.1 Flash can compete with other leading AI models.
It is what DeepSeek has built the architecture to become.
And in an AI race where major laboratories are pushing capability forward at extraordinary speed, the answer could determine whether DeepSeek again forces the rest of the industry to reconsider how much advanced intelligence should cost.

