DeepSeek V4.1 Flash cuts KV cache memory fourfold and pushes V4 Pro into retirement
On September 10, 2026, DeepSeek shipped V4.1 Flash, an open multimodal model that cuts KV cache memory fourfold and lowers API prices by 11 to 57 percent; V4 Pro requests will be routed to Flash starting September 14. Re-benchmark your workloads before that date.