Programmatic tool calling 270 s timeout vs Sume's 55 s script_run cap
Claude times out a pending programmatic tool call after about 4 minutes. Sume's script_run stops at 55 seconds. How to size slow video jobs around both.

Two clocks apply when a model loops over tools: the host's and the tool server's. Anthropic's programmatic tool calling page, read on 2026-10-04, says a pending tool call times out after roughly 4 minutes, surfaced as a 270-second TimeoutError, that idle containers are reclaimed after about 5 minutes, and that a container can live at most 30 days. Sume's script_run is tighter: timeout_seconds runs from 5 to 55, with a default of 45.
Video generation usually takes longer than either clock, so neither loop should wait on a render.
What does each timeout cover?
| Clock | Value | Scope |
|---|---|---|
| Claude pending tool call | about 4 minutes (270 s) | one call waiting for your result |
| Claude idle container | about 5 minutes | sandbox reclaimed |
Sume script_run | 5 to 55 s, default 45 | the whole script |
Sume jobs_wait | up to 55 s per slice, 20 ids | one wait call |
| Sume tool output | 256 KiB serialized | one tool result |
How should I split the work?
Submit in one step, wait in another. A script submits the creates and returns the jobs[] it received, then your agent loop calls jobs_wait in slices of at most 55 seconds, then jobs_result. A slice that returns early with jobs still running is normal; call it again with the ids that remain.
That pattern keeps every Sume call under the 270-second host limit with room to spare, because no single call holds longer than 55 seconds. The Jobs and results page lists the status values and result calls.
A script that submits and returns ids
The sketch below is the body for script_run, passed in the script field. It reads args.scenes as a list of prompt strings and runs generate_image once per scene with its own idempotency key. Check the live tools_schema for the exact generate_image fields before using it, since the field set can change.
const out = [];
for (let i = 0; i < args.scenes.length; i++) {
const r = await sume.tools.generate_image({
prompt: args.scenes[i],
idempotency_key: "scene-" + args.run + "-" + i,
});
out.push(r);
}
return out;What happens on a stop?
If the script hits its time or call budget, the run ends with an error code such as script_timeout, but calls[] and jobs[] are still complete. Never resubmit a create that already registered a job. Reuse the same idempotency_key on any retry, so a repeat returns the original rather than a second charge. The MCP overview lists the other tools.
How do I test the split?
Run a two-job experiment before you build a pipeline. Submit one cheap create and one slow one in a single script, return both job ids, and confirm the script ends well inside its 45-second default. Then call jobs_wait with both ids and watch which slice returns the fast job first. If a slice returns with one job unfinished, call it again with only that id. Record the time of each slice; none should approach the 270-second host limit, and the 55-second wait cap is the number to design around.
Sources
Related posts
More in Developers
- Prometheus counters for Sume API errors, split by code and status
Wrap every Sume call in a counter and histogram labeled by route template, status and error.code, never by job id or request id, then alert on retryable rates.
- Export Sume job counts by status to Prometheus with a Python gauge
A small exporter pages GET /v1/jobs with next_cursor and sets a prometheus_client Gauge labelled by status, so a dashboard shows queued and failed jobs.
- Push or poll for a finished render: listen, webhook or jobs_wait
MCP 2026-07-28 adds subscriptions/listen. For a render that takes minutes, compare a listen stream, a signed webhook and jobs_wait, with a Python verifier.
- Pydantic AI slot leak vs Sume queue_full: tell them apart
Pydantic AI v2.53.0 fixed a streamed-request concurrency slot leak. A client limiter is not Sume's workspace queue_full 429, and each needs its own handling.
Written by Sume