How long can a Sume agent run take? Timeouts for slow models

Format runs are finalized at 90 minutes, or after 10 silent minutes once past 25. Agent Completions publish no deadline. Webhooks time out at 10 s per attempt.

4 min readSume
All posts

Short answer

The Format runs docs give a hard ceiling: a non-terminal receipt carries an expires_at, and Sume force-finalizes the run as failed 90 minutes after created_at, or earlier if the run is older than 25 minutes and has been silent for 10. The Agent Completions page publishes no deadline of its own. It says only that the statuses match Action runs and that you poll status_url until next_action is not poll_status.

Neither the OpenAI page for GPT-6 Astra nor the Anthropic models overview, both read 2026-10-08, states a per-turn timeout, so switching to a faster or slower model does not change these Sume bounds.

The numbers in one place

The model changes how long a turn thinks, not the platform ceiling. Slower reasoning models eat the same 90 minutes.

Timing limits in the Sume docs, read 2026-10-08
LimitValueWhere
Format run hard deadline90 minutes from created_at (expires_at)Format runs, Poll
Early finalizeRun older than 25 minutes and silent for 10Format runs, Poll
Stalled-run signalProgress clock (last timeline at) not moving for several minutesFormat runs, timeline
Long-form video15 to 30 minutes of workFormat runs, Poll
Run webhook attempt10 s per attempt, up to 10 attemptsRun webhooks, Delivery behavior
Agent Completion deadlineNot publishedAgent Completions

How to poll

The docs advise doubling the gap between polls up to one minute, since long-form video takes 15 to 30 minutes and a poll each second wastes read budget. A 429 or 503 while polling does not mean the run failed; it keeps executing and spending. Use expires_at as your own ceiling instead of inventing a number.

For an unattended job, prefer the webhook: it fires once when the run completes or fails.

What a slow model changes

It changes spend more than time. A deeper-thinking model produces more output tokens, and Astra's output is $50 per million versus Haiku 5.5's $0.50. If a run is nearing its cap and the clock, cancel it via POST /v1/agent-runs/{id}/cancel rather than waiting for the force-finalize.

Using the numbers

Two practical conclusions follow. First, size your own client timeout from expires_at, not from the model's speed: a Format run is allowed 90 minutes, so a 2 minute HTTP timeout on the create call is fine only because the create returns immediately with a 202. Second, do not treat a slow poll as a failed run: the docs say a 429 or 503 during the loop is temporary and the run continues to spend.

If you are comparing a slow high-effort model with a fast one, measure wall-clock time to a terminal status on your own Format, not a single turn. A model that thinks longer per turn may take fewer turns, and the platform bounds are the only fixed numbers.

For a concrete plan, imagine a Format run that renders one 10 second clip. The media step is a single Seedance 2 call, and the agent turns around it are a handful of short model calls. Even with a slow model, the run should finish in minutes, well under the 25 minute mark where the silence rule starts to apply. The bounds exist for the unusual run: a stuck tool call, a dead provider or a loop. They are a backstop for the platform, and your client should treat reaching one as a bug worth reporting, not as normal operation.

  • Poll gap: double up to one minute.
  • Cancel with POST /v1/agent-runs/{id}/cancel for an Agent Completion.

Sources

Related posts

More in Models

All Models posts

Written by Sume