DeepSeek introduced V4.1 Flash in September, with native visual understanding and improvements to text and agent tasks. Its official news page dates the release to September 10; the original Chinese post referred to September 9.

The original post highlights a smaller active-parameter footprint and changes to caching that reduce memory and storage demands. It reports HBM requirements reduced to one-quarter and SSD requirements to one-eighth of the previous generation.
Peak and off-peak pricing
The Chinese post quotes off-peak rates of RMB 1 per million uncached input tokens, RMB 4 per million output tokens, and RMB 0.02 per million cached input tokens, with peak prices twice as high.
For U.S. developers, the current international pricing page lists the following dollar rates per million tokens:
| Token type | Off-peak | Peak |
|---|---|---|
| Uncached input | $0.15 | $0.30 |
| Cached input | $0.003 | $0.006 |
| Output | $0.60 | $1.20 |
The documented API model name is deepseek-flash. Check the official pricing page for the peak-hour schedule and the rate applicable to your account.
Sources: DeepSeek announcements and API model details and pricing.
Adapted from the original Chinese article, published on September 12, 2026.

Comments NOTHING