DeepSeek Cuts Flash Model Prices: Output Drops to 4 Yuan per Million Tokens
DeepSeek is making waves again with a significant price reduction for its Flash series models. Starting at noon Beijing time on September 10, 2026, the company will lower costs across the board, making advanced AI more accessible than ever.
If you've been keeping an eye on AI pricing trends, you know that cost is often a barrier for developers and businesses. DeepSeek's latest move aims to tear down that barrier. The Flash series, which includes the deepseek-v4-flash and deepseek-v4-flash-vision-exp models, will see uniform price cuts across all tiers.
Let's break down the numbers. During off-peak hours, the input cost for cache hits drops from 0.05 yuan to just 0.02 yuan per million tokens—a whopping 60% decrease. Peak hours see a similar reduction, from 0.10 yuan to 0.04 yuan. For cache misses, the off-peak price falls from 1.5 yuan to 1 yuan, and peak hours drop from 3.0 yuan to 2 yuan, marking a 33.33% cut. Output costs also decrease, albeit more modestly: off-peak drops from 4.5 yuan to 4 yuan, and peak hours from 9.0 yuan to 8 yuan, an 11.11% reduction.
What does this mean in practice? For developers running large-scale applications, these savings can add up quickly. Imagine processing millions of tokens daily—the reduced output cost alone could translate into significant budget relief. And with cache hits becoming 60% cheaper, repeated queries become far more economical, encouraging more efficient use of the models.
DeepSeek's pricing strategy reflects a broader trend in the AI industry: making powerful models more affordable to foster innovation. By lowering the barrier to entry, they're enabling startups and independent developers to experiment and build without worrying about runaway costs.

It's also worth noting that the price adjustment applies uniformly to both models in the Flash series, ensuring consistency for users who switch between them. Whether you're using the standard flash model or the vision variant, you'll benefit from the same reduced rates.
As AI adoption continues to grow, pricing remains a critical factor. DeepSeek's decision to cut prices, especially for cache hits, signals a commitment to supporting the developer community. It's a smart move that could attract more users to their platform and solidify their position in the competitive AI landscape.
For those already using DeepSeek's Flash models, this is welcome news. For those on the fence, this might be the perfect time to dive in. After all, with prices this low, the cost of experimentation just got a whole lot friendlier.
Key Points
- Effective Date: September 10, 2026, at 12:00 Beijing time.
- Models Affected: deepseek-v4-flash and deepseek-v4-flash-vision-exp.
- Off-Peak Pricing: Input (cache hit) at 0.02 yuan, input (cache miss) at 1 yuan, output at 4 yuan per million tokens.
- Peak Pricing: Double the off-peak rates.
- Biggest Cut: Cache hit prices reduced by 60%.
- Overall Impact: Significant cost savings for developers and businesses using AI at scale.