How long can a Sume agent run take? Timeouts for slow models
Format runs are finalized at 90 minutes, or after 10 silent minutes once past 25. Agent Completions publish no deadline. Webhooks time out at 10 s per attempt.

Short answer
The Format runs docs give a hard ceiling: a non-terminal receipt carries an expires_at, and Sume force-finalizes the run as failed 90 minutes after created_at, or earlier if the run is older than 25 minutes and has been silent for 10. The Agent Completions page publishes no deadline of its own. It says only that the statuses match Action runs and that you poll status_url until next_action is not poll_status.
Neither the OpenAI page for GPT-6 Astra nor the Anthropic models overview, both read 2026-10-08, states a per-turn timeout, so switching to a faster or slower model does not change these Sume bounds.
The numbers in one place
The model changes how long a turn thinks, not the platform ceiling. Slower reasoning models eat the same 90 minutes.
| Limit | Value | Where |
|---|---|---|
| Format run hard deadline | 90 minutes from created_at (expires_at) | Format runs, Poll |
| Early finalize | Run older than 25 minutes and silent for 10 | Format runs, Poll |
| Stalled-run signal | Progress clock (last timeline at) not moving for several minutes | Format runs, timeline |
| Long-form video | 15 to 30 minutes of work | Format runs, Poll |
| Run webhook attempt | 10 s per attempt, up to 10 attempts | Run webhooks, Delivery behavior |
| Agent Completion deadline | Not published | Agent Completions |
How to poll
The docs advise doubling the gap between polls up to one minute, since long-form video takes 15 to 30 minutes and a poll each second wastes read budget. A 429 or 503 while polling does not mean the run failed; it keeps executing and spending. Use expires_at as your own ceiling instead of inventing a number.
For an unattended job, prefer the webhook: it fires once when the run completes or fails.
What a slow model changes
It changes spend more than time. A deeper-thinking model produces more output tokens, and Astra's output is $50 per million versus Haiku 5.5's $0.50. If a run is nearing its cap and the clock, cancel it via POST /v1/agent-runs/{id}/cancel rather than waiting for the force-finalize.
Using the numbers
Two practical conclusions follow. First, size your own client timeout from expires_at, not from the model's speed: a Format run is allowed 90 minutes, so a 2 minute HTTP timeout on the create call is fine only because the create returns immediately with a 202. Second, do not treat a slow poll as a failed run: the docs say a 429 or 503 during the loop is temporary and the run continues to spend.
If you are comparing a slow high-effort model with a fast one, measure wall-clock time to a terminal status on your own Format, not a single turn. A model that thinks longer per turn may take fewer turns, and the platform bounds are the only fixed numbers.
For a concrete plan, imagine a Format run that renders one 10 second clip. The media step is a single Seedance 2 call, and the agent turns around it are a handful of short model calls. Even with a slow model, the run should finish in minutes, well under the 25 minute mark where the silence rule starts to apply. The bounds exist for the unusual run: a stuck tool call, a dead provider or a loop. They are a backstop for the platform, and your client should treat reaching one as a bug worth reporting, not as normal operation.
- Poll gap: double up to one minute.
- Cancel with
POST /v1/agent-runs/{id}/cancelfor an Agent Completion.
Sources
Related posts
More in Models
- How many aspect ratios does each Sume image model list? Oct 2026
Nano Banana 2.1 lists 15 aspect ratio values, Imagen 5. A count for every Sume image row, and why the number decides banners, stories and 4:5 posts.
- How many aspect ratios each Sume image model lists, vs Hy's 14
Hy Image 3.5 Preview lists 14 aspect ratios. On Sume, Nano Banana 2.1 and Ideogram list 15, Grok and Flux 13, GPT Image 2.5 only 9. Full counts by row.
- HunyuanVideo VRAM: 45 GB at 544p, 60 GB at 720p, 80 GB advised
The HunyuanVideo README lists 45 GB for 544x960 and 60 GB for 720x1280 at 129 frames, tested on one 80 GB GPU. FP8 saves about 10 GB. Table and hosted route.
- HunyuanVideo 720p: 1,904 s on one GPU, 338 s on eight
README latency for a 1280x720, 129-frame, 50-step clip: 1,904 s on 1 GPU and 338 s on 8. The GPU-seconds each setup burns, and what it means for hosted jobs.
Written by Sume