Sume job_failed_terminal: retryable true means reuse the same key
On job_failed_terminal never resubmit the identical payload. If error.retryable is true, re-issue the create with the same idempotency_key; else fix the input.

When a Sume MCP job fails, the result carries agent.recovery.code: job_failed_terminal. Do not resubmit the identical payload. If error.retryable is true, re-issue the same create with the same idempotency_key; if it is false, fix the input and ask the user before spending again.
Read the failure from the right tool
A failed job is a terminal state, and jobs_result is for successful jobs: it answers 409 job_not_completed on a failed one. Read the failure with jobs_get, which carries error.public_reason and the retryable flag. The failed job post shows that read; this post is about what to do next.
The decision
The recovery hint reduces the choice to one flag.
| error.retryable | Same payload? | idempotency_key | Next step |
|---|---|---|---|
| true | yes, unchanged | the same key | re-issue the create once, then jobs_wait |
| false | no | a new key once the input changes | fix the input, ask the user, then create |
| absent | no | do not guess | read jobs_get public_reason, then ask |
A small router
Put this in the harness, so the model never decides it.
def after_failure(job: dict, key: str) -> dict:
err = job.get('error') or {}
if err.get('retryable') is True:
return {'action': 'recreate', 'idempotency_key': key}
return {
'action': 'ask_user',
'reason': err.get('public_reason', 'unknown'),
}
print(after_failure({'error': {'retryable': True}}, 'k-1'))
print(after_failure({'error': {'public_reason': 'input_rejected'}}, 'k-1'))Limits
Operations-stopped jobs are a separate case with their own recovery code, and polling them again will not change the answer. A retry is also not a guarantee: if the same failure repeats, stop and report it rather than looping. Tell the user in plain words what failed, what you tried and what it would cost to try again, so the next decision is theirs.
The failure you are avoiding
An agent sees a failed job, concludes it must try again, and builds a fresh request with a new key. If the first job had in fact produced a charge or was still settling, the second create is a second charge. Reusing the key tells Sume the retry is the same intent as the first call. The opposite mistake is resending an identical payload that failed for a reason inside the input, which fails again the same way. The retryable flag exists to separate those two cases.
Checklist before you ship
- Read error.retryable from jobs_get, not from the wording of the message.
- Reuse the original idempotency_key only when retryable is true.
- Cap retries at one per job so a bad payload cannot loop.
- When retryable is false, change the input and ask the user first.
Sources
Related posts
More in Developers
- max_spend_usd on Sume MCP: floors to micros, missing estimate fails
max_spend_usd takes 0 to 10,000, floors to micros, and runs on the preview before dry_run returns. No usable estimate means missing_usage_estimate, not a pass.
- Sume MCP OAuth metadata: curl the well-known URLs before you connect
Before a client fails at sign-in, curl Sume's protected-resource and authorization-server metadata on mcp.sume.com and check the resource audience matches.
- Avatar preview failed on Sume MCP: stop_and_ask, do not generate video
When an avatar video preview fails or is canceled, Sume MCP returns recovery code stop_and_ask: do not call generate_video, report the reason, offer a rewrite.
- timeline_get 409 job_not_completed on Sume MCP: retry, do not give up
A 409 job_not_completed from timeline_get means the render is still running, so retry. Call jobs_wait on the same job_id; never report the run blocked.
Written by Sume