Agent says Sume MCP is down after one timeout: retry once, name it
One errored or timed-out Sume MCP call is just that call failing. Retry it once, report that tool's error, and never say the server is down.

Sume's hosted MCP tells agents, in its initialize instructions, that a tool call which errors or times out is that one call failing: retry it once and report that tool's error. The agent must never describe the MCP or generation server as disconnected, down or still connecting.
What the instruction covers
The rule sits in the run contract the server sends when a client connects. It has three parts: scope the failure to a single call, allow one retry, and report the specific tool and its error. It is paired with a second rule for waits: on timed_out or an HTTP 524, retry the wait or status read on the same ids and do not resubmit the paid create.
Call failure versus server failure
These are different facts and need different words.
| What happened | Say | Do not say |
|---|---|---|
| One tool returned an error | tool name plus its error text | the server is down |
| A wait slice timed out | the job is still running | the job failed |
| A 524 on jobs_wait | a proxy timed out; waiting again | blocked |
| A paid create errored before a job id | retry once with the same key | start a fresh create |
Retry once, with the key
The retry must be safe. For reads it always is. For paid creates, safety comes from the idempotency_key: the retry carries the same one, so a request that did reach Sume returns the same job instead of making a second. The 524 post covers the proxy case in detail.
import time
def call_once_more(fn, *args, **kwargs):
try:
return fn(*args, **kwargs)
except Exception as first:
time.sleep(1)
try:
return fn(*args, **kwargs)
except Exception as second:
name = getattr(fn, '__name__', 'tool')
raise RuntimeError(f'{name} failed twice: {second}') from first
print(call_once_more(lambda: 'ok'))Limits
One retry is the contract for transient errors, not for refusals. An insufficient_scope, insufficient_credits or validation error will not change on a second try; read it and act on it.
Why the wording matters
Agents write what they see. After one timeout, a model that has no rule often tells the user that the server is unreachable and abandons a run that has several healthy calls left. Naming the failure at the level of one call keeps the rest of the run alive and gives the user an accurate sentence to act on, such as which tool failed and what it said.
Checklist before you ship
- Put the retry-once rule in the system prompt, in the same words the server uses.
- Log the failing tool name and its error next to every retry.
- Never switch to a different provider because one call failed.
- Do not retry a paid create without its original idempotency_key.
Sources
Related posts
More in Agents
- Sume MCP tool_name_typo: avatar_image_to_video_create is a typo
avatar_image_to_video_create gives tool_not_found; the real name is avatar-image-to-video_create. Sume sends did_you_mean and a tools_schema next_step.
- Sume MCP returned a tool descriptor, not a result: it did not run
If a Sume MCP call returns a tool name, description and schema instead of a result, the tool did not run. Re-call with the namespaced name your client lists.
- Sume schedule cron field: expr, IANA timezone and next_run_at
A Sume schedule's cron object holds expr, timezone and next_run_at, or null for an API-only schedule. How to read when the next run is due over the API.
- Sume schedule cap null or 0: what a per-run override can do
A per-run generation_spend_cap_usd lowers a Sume schedule's cap but never raises it; null drops that ceiling, 0 is rejected, and wallet limits still apply.
Written by Sume