Summary
On September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash, an MIT-licensed open-weights model with a 1M-token context and a 552B-parameter sparse architecture that activates 8B parameters per input token. It scores 74.2 on DeepSWE v1.1 and launches at $0.15 per million input and $0.60 per million output tokens off-peak, retiring V4 Flash and taking over V4 Pro traffic on September 14.
What changed
DeepSeek shipped V4.1-Flash on its API (model string deepseek-flash) with MIT-licensed weights on Hugging Face. It uses a 552B-parameter backbone with 8B active per input token (16B for output), native image input, a KV cache one quarter the size of V4 Flash, a 1M-token context, and a 74.2 DeepSWE v1.1 score. Off-peak pricing is $0.15/$0.60 per million input/output tokens ($0.30/$1.20 at peak); it retires V4 Flash and absorbs V4 Pro traffic on September 14.
Why it matters
Open MIT weights plus aggressive pricing keep pressure on both closed frontier labs and other open-weight providers, especially on cost-sensitive coding and agent workloads where a 74.2 DeepSWE score and 1M context are competitive. Time-of-day pricing and a smaller KV cache signal a focus on serving economics at scale.
Evidence excerpt
DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026 with MIT weights, 1M context, and a score of 74.2 on DeepSWE v1.1. Pricing is $0.15 per million input tokens and $0.60 per million output off-peak. It retires V4 Flash and takes over V4 Pro traffic on September 14.