DeepSeek Releases V4 Flash and V4 Pro — Open-Source With 1M Context
DeepSeek's latest open-source models claim a 1-million-token context window with improved agentic capabilities and a new Hybrid Attention Architecture. Performance is competitive with US frontier models at dramatically lower cost.
DeepSeek has released two new open-source models, V4 Flash and V4 Pro, with a claimed context window of 1 million tokens. That puts them in a small group of models designed to work with very large inputs, such as long codebases, large document sets, or extended multi-step agent workflows.
The company says the models use a new Hybrid Attention Architecture, a design meant to balance quality and efficiency when handling long context. In practical terms, attention is the mechanism that helps a model decide which parts of the input matter most at each step. Hybrid approaches try to reduce the cost of processing huge prompts while keeping the model responsive across long stretches of text.
⚡ New to this?
This matters because a context window is how much text a model can keep in mind at once, and 1 million tokens is a very large amount. A token is a chunk of text, so this can mean whole code repositories, long documents, or many chat turns in one session. Open-source means the model weights are available to use and modify, which is important for companies that want more control over where their data goes and how the system is deployed.
Agentic capabilities refer to models that can carry out multi-step tasks, not just answer one question at a time. Non-experts should care because these models are increasingly used in business tools, support systems, and automation workflows, and lower cost can affect which teams can actually deploy them.
🦞 OpenClaw angle
Open-source, massive context, strong on agent tasks, and cheap. If you're cost-conscious or want to self-host, DeepSeek V4 is now a serious contender for OpenClaw agents.
DeepSeek is positioning the release as an open-source option for developers and enterprises that want frontier-style capability without the price tag often associated with top proprietary systems. According to the company, V4 Flash and V4 Pro are competitive with leading U.S. models on a range of tasks, including those that rely on agentic behavior, where the model needs to plan, use tools, and carry state across multiple steps.
That agent focus matters because a lot of AI systems are moving beyond single-turn chat. Companies are using models to summarize logs, inspect repositories, draft tickets, query internal tools, and coordinate longer workflows. A model that can retain more context is better suited to those jobs, especially when the input spans many files, pages, or exchanges.
The 1-million-token claim is the headline feature, but the combination of open source and lower cost is what gives the release strategic weight. Open-source models can be downloaded, modified, and run by organizations that want more control over data, deployment, or fine-tuning. That is one reason they are popular in self-hosted environments and private AI stacks.
DeepSeek has built a reputation for releasing models that compete aggressively on price and capability, and this update appears to extend that approach. The company is not alone in pushing longer context windows, but a million tokens is still a notable number because it changes what teams can reasonably keep in memory during a single session.
Long context does not eliminate the need for retrieval systems, databases, or good prompt design. It does, however, reduce how often a model has to be fed information in chunks, which can make agent workflows simpler to build and less brittle in some cases.
The release also reflects a broader shift in the AI market. Model vendors are no longer competing only on benchmark scores, they are also competing on deployment flexibility, context length, and total cost of use. For many buyers, those factors matter more than a narrow win on a test.