ColinBuilds · Current official DeepSeek API pricing

DeepSeek-V4-Flash

Input / 1M
$0.44
Output / 1M
$1.32
Current official context window
1,000,000 tokens

DeepSeek-V4-Flash-0731 / deepseek-v4-flash. Calculator uses peak cache-miss input and peak output. Peak: cache-hit $0.014, cache-miss $0.44, output $1.32 per 1M. Off-peak is half: $0.007 / $0.22 / $0.66. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. · as of August 25, 2026

Compare models

Same per-1M input/output rates as above. Enter your monthly token volumes to compare bills.

DeepSeek-V4-Flash

Input: 1M × $0.44 = $0.44
Output: 1M × $1.32 = $1.32
Total
$1.76

GPT-4o

Input: 1M × $2.50 = $2.50
Output: 1M × $10.00 = $10.00
Total
$12.50

Difference

$10.74

GPT-4o costs $10.74 more than DeepSeek-V4-Flash.

Popular comparisons

37 total featuring DeepSeek-V4-Flash.

Open compare hub →

DeepSeek-V4-Flash is a current official DeepSeek API route with 1M context, 384K maximum output, and peak/off-peak pricing that also splits cache-hit and cache-miss input.

What this model is

DeepSeek-V4-Flash is one of DeepSeek’s current API models listed on the official Models & Pricing page, with model version DeepSeek-V4-Flash-0731 and API ID deepseek-v4-flash. It is not the same thing as older DeepSeek-V3 launch pricing.

DeepSeek also lists deepseek-v4-pro and experimental deepseek-v4-flash-vision-exp. Older deepseek-chat and deepseek-reasoner compatibility names were retired after 2026-07-24.

Pricing notes

DeepSeek lists prices per 1M tokens and now uses peak and off-peak rates. Off-peak rates are half of peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. All other hours are off-peak.

For DeepSeek-V4-Flash, the official table lists:

  • Cache-hit input: $0.007 off-peak / $0.014 peak
  • Cache-miss input: $0.22 off-peak / $0.44 peak
  • Output: $0.66 off-peak / $1.32 peak

The calculator on this page uses the peak cache-miss input price and peak output price as the default comparison, because that is the safer budget when you do not yet know whether prompts will hit cache or land in off-peak hours. Repeated-context or off-peak workloads can be substantially cheaper.

DeepSeek also warns that product prices may vary and recommends checking the pricing page regularly.

Benchmarks and specs

DeepSeek’s official pricing table lists DeepSeek-V4-Flash with a 1M context length and a maximum output of 384K tokens. The same table lists JSON Output, tool calls, Responses API, and Anthropic API support.

DeepSeek does not list parameter counts or benchmark scorecards for this route on the checked pricing page. This page does not list benchmark scores until a named evaluation source and test setup are verified.

Best fit

DeepSeek-V4-Flash is best for cost-sensitive API comparisons, repeated-context workloads where caching may apply, and explanations of why “input price” can mean different things depending on cache hit versus cache miss and peak versus off-peak hours.

Sources