DeepSeek’s pitch for V4.1-Flash is straightforward: more capability with less computation. The new model understands images, supports up to one million tokens of context, and narrowly beats the larger V4-Pro in DeepSeek’s own benchmarks. Its weights are available under the MIT license, giving developers room to customize and host it themselves. The substantial memory requirements, however, will likely put self-hosting out of reach for many businesses.
The bigger news is in the API documentation: starting September 14, requests to the existing V4-Pro endpoint will automatically route to Flash until a new Pro model arrives. That means different model behavior without anyone touching the code, billed at Flash rates: $0.30 per million uncached input tokens and $1.20 per million output tokens during peak hours. Our take: put your workflows through another round of testing. Better benchmark scores are welcome, but they won’t tell you whether your support automation or document processing still works reliably after the switch.