Sume hosted MCP wait_busy 429: retry only the rejected call

wait_busy means the host was full and nothing ran. Retry only that call with the same idempotency_key after retry_after_seconds. Accepted calls keep going.

5 min readSume
All posts

When Sume's hosted MCP server returns wait_busy, the call you sent never started. The error says the MCP host stayed at capacity for this call, that nothing ran, and that you should retry only that rejected call with the same idempotency_key after retry_after_seconds. Calls already accepted keep running. So the right agent behavior is a narrow retry, not resubmitting the whole batch.

This comes from the server's tool admission code and the behavior described in Sume's MCP tools and gates and Jobs and results docs.

What does wait_busy look like?

It returns as an error tool result with code: wait_busy, http_status: 429, retryable: true and retry_after_seconds: 1. The server waits up to 20 seconds for a free slot before it gives up, so the rejection means it really was full for that long.

There is a separate wait_busy for jobs_wait. Its message says the wait budget is full and to keep one jobs_wait for all pending job ids. That one is fixed by merging waits, not by retrying.

Two different wait_busy cases (read 2026-10-03)
ToolWhat it meansWhat to do
Any tool callHost at capacity; nothing ranRetry that call with the same idempotency_key after the delay
jobs_waitWait budget fullUse one jobs_wait with 1 to 20 ids

Why the same idempotency key?

Paid and write tools require an idempotency_key. Reusing the key on the retry means that if the first attempt did land, you do not create a second job. Changing the key turns a safe retry into a duplicate charge. Never mint a new key just because a call was rejected.

What is a safe retry loop?

This helper retries only on wait_busy, waits the stated delay, and keeps the key fixed. It runs offline against a fake caller.

import time

def call_with_retry(call, args, max_tries=4, sleep=time.sleep):
    for attempt in range(max_tries):
        res = call(args)
        err = res.get('error') or {}
        if err.get('code') != 'wait_busy':
            return res
        sleep(err.get('retry_after_seconds', 1))
    raise RuntimeError('still busy after retries')

state = {'n': 0}
def fake(args):
    state['n'] += 1
    if state['n'] < 3:
        return {'error': {'code': 'wait_busy', 'retry_after_seconds': 0}}
    return {'ok': args['idempotency_key']}

print(call_with_retry(fake, {'idempotency_key': 'k-1'}, sleep=lambda s: None))

When is it not a capacity problem?

Do not confuse it with queue_full, which comes from generation admission: that returns 429 too, but it means your plan's generation queue is full, and you slow down at the level of jobs rather than a single call. insufficient_credits is a 402 and a retry will not fix it.

What should an agent not do?

A short rule for an agent prompt: when a tool result is an error with retryable: true, wait the stated delay, retry that single call with the same key, and after the third failure report the call name and key instead of continuing. Calls that were accepted earlier are listed in your own journal, so you never need to ask the server to replay them.

  • Do not resend calls that were already accepted.
  • Do not rotate the key between attempts.
  • Do not poll jobs_wait once per job; one call takes up to 20 ids and caps at 55 seconds a slice.
  • Do not loop forever. Stop after a few tries and report.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume