DeepSeek released DeepSeek-V4.1-Flash on Sept. 10, 2026, pitching the model as the smallest member of a new architecture family with native multimodal vision and lower serving costs than its prior Flash line.

On the DeepSeek API the model is called deepseek-flash. DeepSeek says V4.1 Flash is a 552-billion-parameter mixture-of-experts system with a causal encoder-decoder design that activates about 8 billion parameters on input and 16 billion on output, and that its KV cache needs a fraction of the memory of the previous generation—claims that come from the company, not independent audits. The pricing page lists a 1 million-token context window and up to 384,000 output tokens. Older aliases deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired and temporarily routed to V4.1 Flash.

DeepSeek also cut Flash API prices and kept peak/off-peak rates, with off-peak billed at half of peak. In its launch note the company said third-party tests put V4.1 Flash ahead of V4 Pro on performance, cost and speed and described plans to phase Pro traffic toward Flash. A same-day changelog update, however, said DeepSeek will keep offering V4 Pro API service after Sept. 14, 2026, with billing unchanged—so readers should treat Pro’s long-term status as unsettled even as Flash becomes the default low-cost multimodal SKU.

For rivals, the release is another Hangzhou-lab push to compete on price and throughput in the API tier where OpenAI, Google and Anthropic also sell fast models. Whether V4.1 Flash’s architecture and price cuts hold up outside DeepSeek’s own benchmarks will show up in developer migrations over the next few weeks.