Agent says Sume MCP is down after one timeout: retry once, name it

One errored or timed-out Sume MCP call is just that call failing. Retry it once, report that tool's error, and never say the server is down.

4 min readSume
All posts

Sume's hosted MCP tells agents, in its initialize instructions, that a tool call which errors or times out is that one call failing: retry it once and report that tool's error. The agent must never describe the MCP or generation server as disconnected, down or still connecting.

What the instruction covers

The rule sits in the run contract the server sends when a client connects. It has three parts: scope the failure to a single call, allow one retry, and report the specific tool and its error. It is paired with a second rule for waits: on timed_out or an HTTP 524, retry the wait or status read on the same ids and do not resubmit the paid create.

Call failure versus server failure

These are different facts and need different words.

How to report MCP failures to a user (read 2026-10-05 against the Sume codebase)
What happenedSayDo not say
One tool returned an errortool name plus its error textthe server is down
A wait slice timed outthe job is still runningthe job failed
A 524 on jobs_waita proxy timed out; waiting againblocked
A paid create errored before a job idretry once with the same keystart a fresh create

Retry once, with the key

The retry must be safe. For reads it always is. For paid creates, safety comes from the idempotency_key: the retry carries the same one, so a request that did reach Sume returns the same job instead of making a second. The 524 post covers the proxy case in detail.

import time

def call_once_more(fn, *args, **kwargs):
    try:
        return fn(*args, **kwargs)
    except Exception as first:
        time.sleep(1)
        try:
            return fn(*args, **kwargs)
        except Exception as second:
            name = getattr(fn, '__name__', 'tool')
            raise RuntimeError(f'{name} failed twice: {second}') from first

print(call_once_more(lambda: 'ok'))

Limits

One retry is the contract for transient errors, not for refusals. An insufficient_scope, insufficient_credits or validation error will not change on a second try; read it and act on it.

Why the wording matters

Agents write what they see. After one timeout, a model that has no rule often tells the user that the server is unreachable and abandons a run that has several healthy calls left. Naming the failure at the level of one call keeps the rest of the run alive and gives the user an accurate sentence to act on, such as which tool failed and what it said.

Checklist before you ship

  • Put the retry-once rule in the system prompt, in the same words the server uses.
  • Log the failing tool name and its error next to every retry.
  • Never switch to a different provider because one call failed.
  • Do not retry a paid create without its original idempotency_key.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume