DeepSeek V4.1 Flash peak hours: half price off-peak, one Sume card
DeepSeek bills V4.1 Flash at half price outside 01:00-04:00 and 06:00-10:00 UTC on weekdays. Here is the vendor schedule and what Sume's catalog card says.

DeepSeek charges DeepSeek-V4.1-Flash at two rates: a peak rate during 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, and an off-peak rate that is half of it at every other hour. Off-peak input is $0.15 per 1M tokens and output $0.60; at peak they are $0.30 and $1.20.
Sume's own pricing notes for the model do not follow that clock. They record a single card of $0.30 input and $1.20 output per 1M tokens, and say the earlier scheme that doubled the rate in weekday peak windows is gone. So the vendor's half-price window is the vendor's, and Sume's list does not pass it through.
What is the vendor schedule exactly?
The DeepSeek pricing page lists the Flash model under the API name deepseek-flash, with a 1M context window, 384K maximum output and vision support. Prices are per 1M tokens.
- Peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays.
- Everything else is off-peak, including weekends and Chinese public holidays in full.
| Token type | Off-peak | Peak |
|---|---|---|
| Input, cache hit | $0.003 | $0.006 |
| Input, cache miss | $0.15 | $0.30 |
| Output | $0.60 | $1.20 |
Does the Sume card change by hour?
No. The gateway pricing notes in the Sume repository say DeepSeek V4.1 Flash is $0.30 input and $1.20 output per 1M tokens since 2026-09-28, with a cache read of $0.007, and that the model-level card is the ceiling across the providers the gateway routes to. The peak multiplier from the earlier card is gone.
That makes the Sume number the same as DeepSeek's peak number. A request made at 14:00 UTC on a Wednesday and one made on a Sunday cost the same on Sume's list, so there is no reason to schedule Format runs around the vendor's clock to save on the LLM turn.
Where does a DeepSeek model show up in a Format run?
The model field on a Format run takes an Agents catalog id and selects only the orchestrating LLM; the image, video and audio models are chosen by the Format's tools. An id outside the catalog is a 400 invalid_request, and the receipt echoes the id that ran. The call doc has the field table.
The catalog's retired DeepSeek spellings resolve to V4.1 Flash. Vendor names such as deepseek-flash or deepseek-v4-pro are not catalog ids, so passing them in model is a 400.
What should a team take from this?
If you call DeepSeek directly, the off-peak window can halve the cost of batch work that you can move to nights and weekends in UTC. If you run the model through a Sume Format, price it from Sume's card and read the real figure from the receipt: debited_usd_micros includes the turn's own LLM row, billable_amount_usd_micros does not.
What about the Pro model and legacy names?
The same DeepSeek page lists a second API model, deepseek-v4-pro, at higher prices: cache-hit $0.022 off-peak and $0.044 at peak, cache-miss input $0.66 and $1.32, and output $1.98 and $3.96. It has no vision support. The page also says the older names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted but served by V4.1 Flash at the Flash price.
The rate limits differ too: the page lists a concurrency limit of 2,500 for Flash and 500 for Pro. Those limits are DeepSeek's account limits; they say nothing about how many Format runs your Sume workspace can hold in flight, which the bulk runs doc sets with a concurrency window of 1 to 16.
How should I schedule around it?
For direct DeepSeek use, convert your batch window to UTC and check whether it falls inside the two weekday peak blocks. A job started at 05:00 UTC on a Tuesday is off-peak until 06:00, then peak until 10:00. For Sume runs, the schedule has no effect on the list card, so choose the time that suits your own review, not the vendor's clock.
Sources
Related posts
More in Pricing
- Do more credits make Sume jobs faster? Concurrency is plan-only
No. Sume's processing concurrency is set by plan, not by balance: Free 1, Pro 4, Startup 8, Scale 20. Top-ups only raise what you can reserve.
- Does AI video API billing round up? Sume's ceil rules per endpoint
A 5.2 s clip is billed as 6 s. Which Sume endpoints round seconds or minutes up, which prorate, and worked examples from the published rates.
- Does aspect ratio change AI video price on Sume? No, resolution does
On Sume's per-second video rows the rate is keyed by resolution, not aspect ratio, so 9:16 and 16:9 cost the same. The real price levers, in a table.
- Does Sume use credits or dollars? Reading a credit quote in USD
Sume prices generations in USD per unit and keeps a USD balance; here is how to translate another vendor's credit quote and where to check it.
Written by Sume