DeepSeekSep 10DeepSeek Launches V4.1-Flash Model and Migrates V4-Pro Users
DeepSeek launched DeepSeek-V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model featuring a Causal Encoder-Decoder architecture that activates 8B parameters for input and 16B for output. The model supports a 1-million-token context window and reduces KV cache footprint by 75%. DeepSeek is retiring its V4-Flash variants and will route all V4-Pro API requests to V4.1-Flash starting September 14, 2026.
