A video agent run takes 15 to 30 minutes: design the waiting

Long-form video from an agent is minutes of work, not seconds. Email-me-when-ready, saved drafts and honest limits for a Sume Format run in your product.

4 min readSume
All posts

Plan for a video agent run to take minutes, and design the screen around that. The Sume Format docs say long-form host video usually completes in 15 to 30 minutes, and a run never gives a partial delivery: it either finishes or comes back failed. Your product should let the user leave and come back, not stare at a spinner.

Why this is not a bug

A Format run boots a fresh sandbox, loads the recipe, and lets an agent call generation tools: host takes, B-roll, voiceover, captions and assembly. That is a production, not a model call. Live avatars like Tavus's Griffin-Lite, which the vendor lists at 0.43 seconds average video latency (read 2026-10-07), solve a different problem: a conversation, not a finished cut.

Two timing models (Sume docs and Tavus page, read 2026-10-07)
ModelTypical waitOutput
Sume Format run, long-form host video15 to 30 minutesFinished MP4 on media.sume.com, reviewable
Sume run ceilingForce-finalized as failed at expires_at, at most 90 minutes after creationA failed receipt that still lists artifacts
Griffin-Lite live call0.43 seconds average video latencyA live stream, research preview only

Five patterns that work

  • Return at once: show an order card with the run id the moment you receive the 202.
  • Let users leave: the signed webhook updates your database, so a returning user sees the result.
  • Notify: send an email or in-app message from your webhook handler when the status is completed.
  • Show real phases: queued, preparing, running, finalizing, as the events timeline reports them.
  • Offer cancel: POST /v1/format-runs/{run_id}/cancel stops a run the user no longer wants.

Say what you do not know

Do not show a countdown. You do not know when the run will end, only that it will before expires_at. A message such as Usually 15 to 30 minutes is accurate and costs you nothing. If a run waits in queued longer than the normal window, the receipt says so through queue.state, and you can show that explicitly.

Plan for failure as a normal state

A run that could not finish comes back failed with an error, and its artifacts[] still lists any media it made. Give users a retry that reuses the same input and a new idempotency key, and show what was made. For a single weak scene, continue the thread with previous_run_id instead of paying for the whole run again.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume