DeepSeek Flash or V4 Pro for a Sume job-polling loop: a price split
On DeepSeek's page Flash costs under a third of V4 Pro per token. Use it for jobs_wait polling; keep payload writing on one model. Read 2026-10-05.

On DeepSeek's pricing page, V4.1 Flash costs well under a third of V4 Pro per token on every line (cache miss input is $0.15 against $0.66 per 1M off-peak). For a Sume job-polling loop, where each step only reads outcome and calls jobs_wait again, the cheaper model is usually enough; keep payload writing on one model.
What DeepSeek lists
The pricing page (read 2026-10-05) shows both models with 1M context, up to 384K output, JSON output and tool calls, with Responses API and Anthropic API support. Vision input is listed for Flash only. The model ids are on the docs home: deepseek-flash and deepseek-v4-pro. The older names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted but served by V4.1 Flash.
| Item | deepseek-flash | deepseek-v4-pro | Pro / Flash |
|---|---|---|---|
| Input, cache hit | $0.003 | $0.022 | about 7x |
| Input, cache miss | $0.15 | $0.66 | 4.4x |
| Output | $0.60 | $1.98 | 3.3x |
| Vision input | listed | not listed | - |
Where each step lands in a Sume run
A Sume media run has a few kinds of steps. Writing the payload for a paid create is a creative step with a cost: a poor prompt wastes a generation. Waiting is mechanical: jobs_wait returns an outcome, and the create result's agent.next_step already names the call. Reading a failure through jobs_get and choosing between retry and ask is a rule-following step, and the recovery code gives the rule.
A routing table in code
A harness can pick the model by step kind. The step names are yours.
ROUTE = {
'plan': 'deepseek-v4-pro',
'write_payload': 'deepseek-v4-pro',
'wait': 'deepseek-flash',
'read_failure': 'deepseek-flash',
}
def model_for(step: str) -> str:
if step not in ROUTE:
raise KeyError(f'unknown step {step!r}')
return ROUTE[step]
print(model_for('wait'))Limits
This is a cost judgment, not a quality benchmark; DeepSeek's page gives prices and capabilities, not accuracy on your task. Test your own loop. Whatever model you pick, send max_spend_usd and an idempotency_key on each paid call; see the jobs_wait post for the wait rules.
A practical split
A polling step has a small decision: is the outcome terminal, and if not, wait again. That is a good job for the cheaper model, with the harness following agent.next_step for the arguments. Planning a scene list or writing a payload is where a stronger model may earn its cost. You can run both in one run, as long as the idempotency key and the payload are produced once and passed through unchanged.
Checklist before you ship
- Use the cheaper model for wait and status steps, which carry no decision.
- Keep the model that writes prompts and payloads consistent across a run.
- Remember vision input is on Flash only, per DeepSeek's page.
- Re-check prices and ids on DeepSeek's page before each budget review.
Sources
Related posts
More in Comparisons
- Draft vs final render cost: a 6-second clip on five Sume video models
The final render costs 2.5x to 10x the draft on Sume, depending on the model. A 6-second clip priced at the cheapest and the top tier, per model.
- Eleven v4 tags like [light rain] vs a separate sound bed on Sume
Eleven v4 puts effects such as [light rain] inside the speech. Sume keeps voice and bed separate, with gain_db, loop and duck_db. The trade-off for editing.
- Directing delivery: Eleven v4 audio tags vs Gemini TTS style
Eleven v4 puts direction inline as audio tags like [laughs]; Gemini TTS adds a separate style field. Sume sends transcript, voice and language only.
- Eleven v4 IPA support vs Sume's pronunciation_dict_id for brand names
ElevenLabs says Eleven v4 improves IPA phoneme support. Sume TTS has an optional pronunciation_dict_id. How each helps with brand names today.
Written by Sume