DeepSeek V4.1 Flash peak pricing vs Sume spend caps for agent loops

DeepSeek bills Flash at double rates 01:00-04:00 and 06:00-10:00 UTC on weekdays. Sume spend caps cover generation only, so budget planner tokens separately.

4 min readSume
All posts

DeepSeek's pricing page lists two rates for V4.1 Flash: a peak rate that is twice the off-peak rate, applying 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, excluding Chinese public holidays. Sume's spend caps, max_spend_usd and generation_spend_cap_usd, only bound Sume generation, so an agent loop has two meters to watch.

The DeepSeek side

Per DeepSeek's pricing page (read 2026-10-05), prices are per 1M tokens and split into cache hit, cache miss and output. The model ids on the API docs home are deepseek-flash (V4.1 Flash) and deepseek-v4-pro.

DeepSeek price per 1M tokens, USD, off-peak and peak (DeepSeek pricing page, read 2026-10-05)
ItemFlash off-peakFlash peakPro off-peakPro peak
Input, cache hit$0.003$0.006$0.022$0.044
Input, cache miss$0.15$0.30$0.66$1.32
Output$0.60$1.20$1.98$3.96

The Sume side

When a model you host calls Sume's hosted MCP, Sume charges for the generations the tools start; the planner provider charges for the tokens the loop uses. Different invoices, different caps. For Agent Completions, generation_spend_cap_usd is required on POST /v1/agent/completions and a request without it gets a 400. That cap is Sume's generation budget for the run, not a token budget for any outside model.

Estimate a loop yourself

Plug in your own token counts from your provider's usage fields. The numbers below are arbitrary inputs, not measurements.

RATES = {  # USD per 1M tokens: (miss, output)
    ('flash', 'off'): (0.15, 0.60), ('flash', 'peak'): (0.30, 1.20),
    ('pro', 'off'): (0.66, 1.98), ('pro', 'peak'): (1.32, 3.96),
}

def loop_cost(model, window, in_tokens, out_tokens):
    miss, out = RATES[(model, window)]
    return in_tokens / 1e6 * miss + out_tokens / 1e6 * out

for w in ('off', 'peak'):
    print(w, round(loop_cost('flash', w, 200_000, 20_000), 4))

Limits

The sketch ignores cache-hit tokens, so it overstates a loop that reuses a long prefix. Peak and off-peak depend on UTC and on Chinese holidays, which this post does not list; read the vendor page for the current rule. See also the DeepSeek tool-call post for wiring.

What the cap does not cover

max_spend_usd and an Agent Completion's generation_spend_cap_usd bound Sume generation spend. They say nothing about the tokens your planner model burns while it thinks, calls tools and polls. If an agent loops on jobs_wait all night, the generation bill may be tiny while the planner bill is not. Put a step or token limit in the client, and read the planner provider's invoice separately.

Checklist before you ship

  • Track planner token spend and Sume generation spend as two separate numbers.
  • Schedule heavy planning loops outside the peak window if latency allows.
  • Send max_spend_usd on paid Sume calls whatever the planner costs.
  • Re-read the DeepSeek pricing page before you publish a budget.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume