DeepSeek Opens V4.1-Flash Under MIT as API Reroutes V4-Pro Traffic on September 14

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026 as a 552B MoE multimodal open-weight model under MIT, and starting 04:00 UTC on September 14 its API routes deepseek-v4-pro requests to Flash at Flash rates until V4.1-Pro ships.

DeepSeek V4.1-Flash open MIT multimodal MoE model
DeepSeek V4.1-Flash open MIT multimodal MoE model

DeepSeek published DeepSeek-V4.1-Flash on September 10, 2026 as the smallest model in its new architecture family, with open weights on Hugging Face under the MIT License. The company and SiliconANGLE report a 552-billion-parameter mixture-of-experts multimodal model that natively handles images and text, supports contexts up to one million tokens, and is live on DeepSeek’s API under the deepseek-flash name.

The same announcement sets a hard cutover for older flagship traffic. Starting at 04:00 UTC on September 14, 2026, requests to deepseek-v4-pro are answered by V4.1-Flash and billed at V4.1-Flash rates until a future V4.1-Pro launch. Retired deepseek-v4-flash and deepseek-v4-flash-vision-exp names temporarily route to V4.1-Flash for compatibility.

Architecture and efficiency claims

DeepSeek describes a Causal Encoder–Decoder stack that activates about 8 billion parameters while processing a prompt and about 16 billion while generating output. The Hugging Face model card and technical summary put the global KV cache near 890 bytes per token—roughly one-quarter of V4-Flash’s HBM footprint—and say persistent SSD cache storage falls to about one-eighth of the prior generation. DeepSeek frames those savings as especially relevant for input-heavy agent workloads where cache-hit charges dominate bills.

Benchmarks and open weights

DeepSeek’s own maximum-effort table, also cited by SiliconANGLE, reports 90.6 on Terminal-Bench 2.1 and 74.2% resolved on DeepSWE v1.1, ahead of DeepSeek-V4-Pro on those agentic rows in the vendor comparison; treat the figures as vendor-reported. Weights sit at deepseek-ai/DeepSeek-V4.1-Flash under MIT, with a technical report PDF in the same repository. DeepSeek says it will work with the open-source community on inference support and explore further deployment options.

This brief is about the V4.1-Flash model and API cutover, not the earlier DeepSeek Harness sandbox CVE.

Primary sources are DeepSeek’s September 10 product news post, the Hugging Face model card for DeepSeek-V4.1-Flash, SiliconANGLE’s September 10 launch report, and Analytics India Magazine’s same-day summary.

Topics
  • #Opensource
  • #Products
Raj M

Author

Raj M

Contributor

AI Systems Architect is a seasoned technology leader with over 15 years of experience in the IT industry working with Fortune 500 companies. With a solid foundation in multi-agent systems, open-source LLM infrastructure, and enterprise deployment, he excels at building scalable production-grade AI platforms.