Two model fields: deepseek-flash for your planner, sume-agent for Sume
Your planner config says deepseek-flash. The Sume Agent Completion model field takes only sume-agent. Keep them in separate settings.

Use two separate settings: one for the model that plans, one for the Sume Agent. On DeepSeek's September 10 notice, V4.1 Flash has the API model id deepseek-flash, and the older ids deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to it (read 2026-10-05). On a Sume Agent Completion, model accepts only sume-agent. A single shared model variable would send the wrong id to one of them.
What each field accepts
The Sume docs state it plainly: model is optional on POST /v1/agent/completions, only sume-agent is accepted, and omitting it gives the same agent. Any other value is an invalid_request 400, listed in the errors table alongside a missing spend cap.
| Side | Field | Accepts |
|---|---|---|
| DeepSeek API | model | deepseek-flash for V4.1 Flash |
| DeepSeek legacy ids | deepseek-v4-flash, deepseek-v4-flash-vision-exp | Temporarily route to V4.1 Flash, per the notice |
| Sume Agent Completion | model | sume-agent only, or omit |
| Sume Agent Completion | generation_spend_cap_usd | Required, no default |
The failure this prevents
Agent frameworks often keep one MODEL constant and pass it everywhere. When the planner moves to a new id, a tool wrapper that forwards the constant to Sume turns into a stream of 400 errors. The reverse is quieter: a wrapper that hard-codes sume-agent for the planner would get a model-not-found from the planner API.
Do not forward a model name to Sume at all. Leaving model out is documented to give the same agent.
A wrapper that cannot leak the planner id
The function below builds the request body from only the fields Sume needs. It runs offline and has no model parameter to misuse.
import json
def sume_completion_body(task: str, cap_usd: float) -> dict:
if cap_usd <= 0:
raise ValueError("cap_usd must be positive")
return {"instruction": task,
"generation_spend_cap_usd": cap_usd}
PLANNER_MODEL = "deepseek-flash"
if __name__ == "__main__":
body = sume_completion_body("Cut a 15 second teaser", 2)
assert "model" not in body
print(PLANNER_MODEL, json.dumps(body))What the Sume receipt echoes back
The 202 receipt carries model: sume-agent and a thread_id. Each completion runs in a new thread, and the receipt's thread id identifies it. Log both next to your planner's model id, so a later audit can tell which planner version asked for which run.
If you keep a per-run record, add the planner id, the Sume run id and the cap you sent. That three-field record is enough to reconstruct why a run happened and what it was allowed to spend, without storing prompts or keys.
Keep in mind that Sume treats input as data and not as instructions, so nothing you put in the planner's own config can change what the Sume Agent is told to do. The instruction or messages you send are the only prompt it follows.
When the planner changes again
Model ids on the vendor side will keep moving. The notice already says V4-Pro requests will route to V4.1-Flash from September 14, 2026 (read 2026-10-05). Keep that churn in the planner's config file. The Sume wrapper should not need a release.
- Store the planner id in one config key and read it only in the planner client.
- Leave
modelout of Sume calls. - Add a unit test that the Sume body has exactly
instructionormessagesplus the cap.
Sources
Related posts
More in Developers
- DeepSeek V4.1 Flash peak pricing vs Sume spend caps for agent loops
DeepSeek bills Flash at double rates 01:00-04:00 and 06:00-10:00 UTC on weekdays. Sume spend caps cover generation only, so budget planner tokens separately.
- DeepSeek V4.1 Flash JSON Output vs Sume output_schema: what differs
DeepSeek's JSON Output makes valid JSON but does not enforce a schema. Sume's output_schema is enforced on a run, with output null and output_error on a miss.
- Deno --allow-net=api.sume.com is not enough to save a finished video
Sume's content URL answers 302 to a file host. A Deno script with a narrow --allow-net fails at the second fetch: read the redirect host, then allow it.
- Derive a Sume idempotency key from tenant, order, Format slug, version
Hash stable ids, not a random UUID, so a retry reuses the key. A failed run needs a new key, and a changed body with the same key returns 409.
Written by Sume