Sume SDK wait timeouts: 20 min, 10 min, and the 90-minute run

subscribeFormatRun waits 20 minutes, waitForRun 10, waitForJob 20, yet a run lives up to 90. Which clock fires first and how to resume after a timeout.

6 min readSume
All posts

Three SDK helpers wait for three different lengths of time by default: subscribeFormatRun for 20 minutes, waitForRun for 10, and waitForJob for 20. None of them is the run's own clock, because a Format run can live up to 90 minutes before Sume force-finalizes it. So a wait that ends in SumeRunTimeoutError or SumeJobTimeoutError does not mean the work failed; it means your client stopped looking. The SDK runs page says it directly: a timeout does not cancel the run, and the run keeps going and keeps billing.

The clocks side by side

The 20-minute default on subscribeFormatRun is chosen because video Formats routinely take 10 to 20 minutes, and waitForJob has the same 20 for video and avatar-video jobs. waitForRun is the generic poller for Format, Action and Agent runs, requires a family argument, and defaults to 10. The poll interval is 2 seconds in all three, with one difference: for jobs it is a floor, and next_poll_after_seconds from the status payload wins when it asks for a longer gap.

Two safeguards sit underneath. The helpers tolerate transient read failures, so a brief 429 or 503 on a poll does not end the wait. And an AbortSignal passed as signal aborts the wait and the in-flight request, which is how you enforce a deadline of your own.

read 2026-10-03
ClockDefault lengthPoll gapOn expiry
subscribeFormatRun20 minutes2 sSumeRunTimeoutError
waitForRun10 minutes2 sSumeRunTimeoutError, family required
waitForJob20 minutes2 s floor, next_poll_after_seconds winsSumeJobTimeoutError
The run itselfup to 90 minutesnot applicableforce-finalized by Sume

Resuming instead of resubmitting

The classic mistake is to treat the timeout as a failure and submit again. That creates a second paid run for the same task. Do the opposite: catch the timeout error, record the run id and last status, and keep waiting or read the run later. The SumeRunTimeoutError carries runId and lastStatus for exactly this.

The function below loops over waitForRun until the run's own 90-minute ceiling plus a minute of slack. It needs @sume-com/sdk 0.2.0 and an API key, and markPending is your own persistence hook.

import { createSumeClient, waitForRun, SumeRunTimeoutError } from "@sume-com/sdk";

const client = createSumeClient({ apiKey: process.env.SUME_API_KEY });
const RUN_CEILING_MS = 90 * 60_000; // the run's own expires_at ceiling

export async function watch(runId, startedAt, markPending) {
  while (Date.now() - startedAt < RUN_CEILING_MS + 60_000) {
    try {
      // Default timeout here is 10 minutes; ask for what this Format needs.
      return await waitForRun(runId, { client, family: "format", timeout: 20 * 60_000 });
    } catch (error) {
      if (!(error instanceof SumeRunTimeoutError)) throw error;
      await markPending(error.runId, error.lastStatus); // still running, still billing
    }
  }
  throw new Error(`run ${runId} should have been force-finalized by now; read its receipt`);
}

Choose a deadline per Format

Set the wait from what the Format does, not from the default. A still-image Format that finishes in a minute should time out at three, so a stuck run is visible quickly. A video Format with several generations should get the full 20 minutes, and a long assembly may justify the loop above. Keep the deadline in the Format's config next to its spend cap.

If your process cannot stay alive that long, stop polling and use a run webhook instead, with a poll as the fallback.

What a hit at 90 minutes means

expires_at is a force-finalize ceiling, or sooner when the run goes silent. A run that reaches it is read as a terminal receipt, not a hang. Read the status and error rather than assuming success, since terminal is not success, and use previous_run_id to continue if the work was incomplete.

Bulk and fleets

If you wait on many runs from one process, remember that polling is one timer and one open request per run in flight, and the helpers jitter their polls so clients started together do not stay in phase. Past a few dozen concurrent runs, prefer webhooks and a sweeper, and keep the SDK wait for the cases that need an answer inline. Cancel explicitly with cancelFormatRun for runs you no longer want, since stopping the wait does nothing to the spend.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume