DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient
π Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
πΉ Introducing the smallest model in our new architecture family, with native visual understanding.
πΉ Designed for greater capability, faster inference, higher throughput, and scaling to larger models.

π§ Asymmetric architecture. More intelligence, less cost.β
πΉ 552B-parameter MoE.
πΉ New Causal EncoderβDecoder architecture: just 8B active parameters for input, 16B for output.
πΉ New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

πΎ Smaller KV cache. Bigger savings.β
Compared with the previous generation, V4.1-Flash's KV cache needs just:
πΉ 1/4 the HBM
πΉ 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.

β‘ V4.1-Flash is now live on the DeepSeek API with native multimodal support.β
Set your model to deepseek-flash.
πΉ V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
πΉ Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. We're phasing out V4-Pro.
πΉ Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.
π€ Official partners WorkBuddy (including CodeBuddy) & OpenCode now fully support V4.1-Flash. Try it today!
π° More efficient architecture. Lower API prices.β
V4.1-Flash lets us serve more users at a lower cost. We're passing the savings on to you.
πΉ Peak/off-peak pricing continues to balance demand.
πΉ Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
πΉ New pricing takes effect at 04:00 UTC on Sept 10, 2026.

π Supporting open source. Expanding deployment options.β
We'll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let's talk.
πΉ Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
πΉ Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf