DeepSeek V4.1 Flash peak pricing vs Sume spend caps for agent loops
DeepSeek bills Flash at double rates 01:00-04:00 and 06:00-10:00 UTC on weekdays. Sume spend caps cover generation only, so budget planner tokens separately.

DeepSeek's pricing page lists two rates for V4.1 Flash: a peak rate that is twice the off-peak rate, applying 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, excluding Chinese public holidays. Sume's spend caps, max_spend_usd and generation_spend_cap_usd, only bound Sume generation, so an agent loop has two meters to watch.
The DeepSeek side
Per DeepSeek's pricing page (read 2026-10-05), prices are per 1M tokens and split into cache hit, cache miss and output. The model ids on the API docs home are deepseek-flash (V4.1 Flash) and deepseek-v4-pro.
| Item | Flash off-peak | Flash peak | Pro off-peak | Pro peak |
|---|---|---|---|---|
| Input, cache hit | $0.003 | $0.006 | $0.022 | $0.044 |
| Input, cache miss | $0.15 | $0.30 | $0.66 | $1.32 |
| Output | $0.60 | $1.20 | $1.98 | $3.96 |
The Sume side
When a model you host calls Sume's hosted MCP, Sume charges for the generations the tools start; the planner provider charges for the tokens the loop uses. Different invoices, different caps. For Agent Completions, generation_spend_cap_usd is required on POST /v1/agent/completions and a request without it gets a 400. That cap is Sume's generation budget for the run, not a token budget for any outside model.
Estimate a loop yourself
Plug in your own token counts from your provider's usage fields. The numbers below are arbitrary inputs, not measurements.
RATES = { # USD per 1M tokens: (miss, output)
('flash', 'off'): (0.15, 0.60), ('flash', 'peak'): (0.30, 1.20),
('pro', 'off'): (0.66, 1.98), ('pro', 'peak'): (1.32, 3.96),
}
def loop_cost(model, window, in_tokens, out_tokens):
miss, out = RATES[(model, window)]
return in_tokens / 1e6 * miss + out_tokens / 1e6 * out
for w in ('off', 'peak'):
print(w, round(loop_cost('flash', w, 200_000, 20_000), 4))Limits
The sketch ignores cache-hit tokens, so it overstates a loop that reuses a long prefix. Peak and off-peak depend on UTC and on Chinese holidays, which this post does not list; read the vendor page for the current rule. See also the DeepSeek tool-call post for wiring.
What the cap does not cover
max_spend_usd and an Agent Completion's generation_spend_cap_usd bound Sume generation spend. They say nothing about the tokens your planner model burns while it thinks, calls tools and polls. If an agent loops on jobs_wait all night, the generation bill may be tiny while the planner bill is not. Put a step or token limit in the client, and read the planner provider's invoice separately.
Checklist before you ship
- Track planner token spend and Sume generation spend as two separate numbers.
- Schedule heavy planning loops outside the peak window if latency allows.
- Send max_spend_usd on paid Sume calls whatever the planner costs.
- Re-read the DeepSeek pricing page before you publish a budget.
Sources
Related posts
More in Developers
- DeepSeek V4.1 Flash JSON Output vs Sume output_schema: what differs
DeepSeek's JSON Output makes valid JSON but does not enforce a schema. Sume's output_schema is enforced on a run, with output null and output_error on a miss.
- Deno --allow-net=api.sume.com is not enough to save a finished video
Sume's content URL answers 302 to a file host. A Deno script with a narrow --allow-net fails at the second fetch: read the redirect host, then allow it.
- Derive a Sume idempotency key from tenant, order, Format slug, version
Hash stable ids, not a random UUID, so a retry reuses the key. A failed run needs a new key, and a changed body with the same key returns 409.
- Hash the avatar payload into your Idempotency-Key, adopt the 409 job
Hash the avatar request into the Idempotency-Key: a retry returns the same job (idempotency_hit true); a 409 names the job holding the key.
Written by Sume