Skip to main content

DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient

πŸš€ Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.

πŸ”Ή Introducing the smallest model in our new architecture family, with native visual understanding.

πŸ”Ή Designed for greater capability, faster inference, higher throughput, and scaling to larger models.


🧠 Asymmetric architecture. More intelligence, less cost.​

πŸ”Ή 552B-parameter MoE.

πŸ”Ή New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.

πŸ”Ή New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.


πŸ’Ύ Smaller KV cache. Bigger savings.​

Compared with the previous generation, V4.1-Flash's KV cache needs just:

πŸ”Ή 1/4 the HBM

πŸ”Ή 1/8 the SSD storage

Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.


⚑ V4.1-Flash is now live on the DeepSeek API with native multimodal support.​

Set your model to deepseek-flash.

πŸ”Ή V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.

πŸ”Ή Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. We're phasing out V4-Pro.

πŸ”Ή Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.

🀝 Official partners WorkBuddy (including CodeBuddy) & OpenCode now fully support V4.1-Flash. Try it today!


πŸ’° More efficient architecture. Lower API prices.​

V4.1-Flash lets us serve more users at a lower cost. We're passing the savings on to you.

πŸ”Ή Peak/off-peak pricing continues to balance demand.

πŸ”Ή Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.

πŸ”Ή New pricing takes effect at 04:00 UTC on Sept 10, 2026.


🌐 Supporting open source. Expanding deployment options.​

We'll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.

Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let's talk.

πŸ”Ή Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

πŸ”Ή Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf