DeepSeek Flash or V4 Pro for a Sume job-polling loop: a price split

On DeepSeek's page Flash costs under a third of V4 Pro per token. Use it for jobs_wait polling; keep payload writing on one model. Read 2026-10-05.

4 min readSume
All posts

On DeepSeek's pricing page, V4.1 Flash costs well under a third of V4 Pro per token on every line (cache miss input is $0.15 against $0.66 per 1M off-peak). For a Sume job-polling loop, where each step only reads outcome and calls jobs_wait again, the cheaper model is usually enough; keep payload writing on one model.

What DeepSeek lists

The pricing page (read 2026-10-05) shows both models with 1M context, up to 384K output, JSON output and tool calls, with Responses API and Anthropic API support. Vision input is listed for Flash only. The model ids are on the docs home: deepseek-flash and deepseek-v4-pro. The older names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted but served by V4.1 Flash.

DeepSeek Flash vs Pro, off-peak USD per 1M tokens (DeepSeek, read 2026-10-05)
Itemdeepseek-flashdeepseek-v4-proPro / Flash
Input, cache hit$0.003$0.022about 7x
Input, cache miss$0.15$0.664.4x
Output$0.60$1.983.3x
Vision inputlistednot listed-

Where each step lands in a Sume run

A Sume media run has a few kinds of steps. Writing the payload for a paid create is a creative step with a cost: a poor prompt wastes a generation. Waiting is mechanical: jobs_wait returns an outcome, and the create result's agent.next_step already names the call. Reading a failure through jobs_get and choosing between retry and ask is a rule-following step, and the recovery code gives the rule.

A routing table in code

A harness can pick the model by step kind. The step names are yours.

ROUTE = {
    'plan': 'deepseek-v4-pro',
    'write_payload': 'deepseek-v4-pro',
    'wait': 'deepseek-flash',
    'read_failure': 'deepseek-flash',
}

def model_for(step: str) -> str:
    if step not in ROUTE:
        raise KeyError(f'unknown step {step!r}')
    return ROUTE[step]

print(model_for('wait'))

Limits

This is a cost judgment, not a quality benchmark; DeepSeek's page gives prices and capabilities, not accuracy on your task. Test your own loop. Whatever model you pick, send max_spend_usd and an idempotency_key on each paid call; see the jobs_wait post for the wait rules.

A practical split

A polling step has a small decision: is the outcome terminal, and if not, wait again. That is a good job for the cheaper model, with the harness following agent.next_step for the arguments. Planning a scene list or writing a payload is where a stronger model may earn its cost. You can run both in one run, as long as the idempotency key and the payload are produced once and passed through unchanged.

Checklist before you ship

  • Use the cheaper model for wait and status steps, which carry no decision.
  • Keep the model that writes prompts and payloads consistent across a run.
  • Remember vision input is on Flash only, per DeepSeek's page.
  • Re-check prices and ids on DeepSeek's page before each budget review.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume