Which Sume audio calls can sync-wait 30 s? A map vs 150 ms claims

Detach, timeline audio, ingest, music and TTS: which can return in a 30 s sync wait, which are async jobs, and what a 150 ms end-to-end claim leaves out.

5 min readSume
All posts

On Sume, no audio call is a streaming socket. Each one creates a job, and mode: "sync" only lets the HTTP request block for up to 30 seconds before it answers with the current job state. A vendor claim such as MAI-Voice-2.1-Flash at 150 ms end-to-end latency (reported by a third-party tracker, not the vendor page, read 2026-10-05) describes a live voice path. It does not describe a job that mints a file you will drop into a video timeline. This page maps which Sume audio calls can finish inside the wait, and which you should poll.

The map

The two audio-only ffmpeg calls (detach and timeline audio) use no provider inference, only worker ffmpeg, and cost $0.01 per job. They are the likeliest to land inside a 30-second wait, but the docs do not promise it. Read the status, not the clock.

Default mode and wait behavior of Sume audio and media calls (read 2026-10-05)
CallDefault modeSync waitWhere it is documented
POST /v1/audio-detachasyncmode sync waits up to 30 s for a 200, else 202Audio detach
POST /v1/timeline-1.0/audioasyncmode sync waits up to 30 s, else 202Timeline audio
POST /v1/timeline-1.0/renderasyncmode sync waits up to 30 s, else 202Timeline 1.0
POST /v1/reference-ingestsyncwaits up to 30 s, else 202 with a queued jobReference ingest
POST /v1/music-router/generatemodes: async, sync, subscribe, webhookwait_timeout_seconds 0 to 30Music 1.0, Music Router
POST /v1/tts-1.0/generate and POST /v1/stt-1.0/transcribejob; poll or webhookNot stated on a Sume docs page I read; check the API referenceJobs and results

What 202 means

A 202 is not a failure. It means Sume accepted the job and it is queued or processing. The jobs and results page says a local timeout is not a reason to submit the same paid request again: store the job id, poll GET /v1/jobs/:id/status with backoff, and read GET /v1/jobs/:id/result once it is completed.

Webhook delivery is terminal-only: job.completed, job.failed and job.canceled. There are no partial callbacks and no progress events. If you wanted partial transcripts while someone talks, that is a different product shape from Sume STT, which returns text and word timings when the job is done.

Why 150 ms and a job do not compare

A first-audio latency number measures how soon sound starts. A Sume voiceover job measures how soon a finished, hosted file exists, with word timings and optional sentence slices. For a short video, you need the whole file before the timeline can place scenes on the spine, so the first byte matters less than the finished artifact.

A live agent that answers a caller needs the first. A 30-second explainer that gets rendered once needs the second. Choose by what your product waits for.

A rule you can apply

  • Short, single-step calls: try mode: "sync" and handle both 200 and 202.
  • Anything chained, such as TTS, then timeline, then captions: submit async, store job ids, and use a webhook plus polling as a backup.
  • Never resubmit a paid job because a client timeout fired.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume