ChengRang

DeepSeek V4.1-Flash

AI Platforms Open Source

DeepSeek's smallest new-architecture model (Sep 10, 2026): 552B MoE with 8B/16B active params, native vision, 1M context, FP4 KV cache at ~890 bytes/token, claimed to outperform V4 Pro across the board; MIT-licensed with API input from 1 yuan per million tokens

LLMOpen SourceMultimodal1M ContextAgents
Visit DeepSeek V4.1-Flash

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

DeepSeek-V4.1-Flash is the smallest model in DeepSeek's brand-new architecture series, released on September 10, 2026. Built on a novel Causal Encoder-Decoder architecture, it is designed for a higher capability ceiling, faster inference and greater throughput, with a path to scale to larger parameter counts.

The model is a 552B MoE that activates about 8B parameters for prefill and 16B for decoding, supporting 1M context with up to 384K output. Its most notable innovation is KV cache compression: the FP4-quantized global KV cache drops to roughly 890 bytes per token, about a quarter of the previous V4-Flash, dramatically lowering the memory and storage cost of long-context agent workloads.

It is also DeepSeek's first official model with native multimodal vision understanding, carrying forward the direction validated by the August Vision-Exp experimental release. The API model name is deepseek-flash, weights are open-sourced under MIT on Hugging Face, and third-party cloud platforms have onboarded it in parallel.

Key Features

Use Cases

Pros

Pricing

Peak/off-peak API pricing (per million tokens): cache-hit input 0.02 yuan off-peak and 0.04 yuan peak; cache-miss input 1 yuan off-peak and 2 yuan peak; output 4 yuan off-peak and 8 yuan peak. Peak hours are weekdays 9:00-12:00 and 14:00-18:00 Beijing time. Concurrency limit 2500. Weights are MIT-licensed with no fee for self-hosting.

Summary

DeepSeek-V4.1-Flash is the most important open-source model update from China in 2026. Its new architecture delivers higher capability and a dramatic cost reduction at the same time, and DeepSeek is confident enough to retire its own flagship V4 Pro in its favor. For teams needing long-context agents, native vision and low-cost scale-out, it is a top pick in the open-source field; for API users, peak/off-peak pricing plus cache discounts push costs to the lowest tier in the industry.

Version History

Category
AI Platforms
Pricing
Open Source
Tags
LLM · Open Source · Multimodal

Related Tools