When is async TTS the right choice? Sync wait, poll or webhook
Async TTS is right for voiceovers, batches and anything a person is not watching a spinner for. Sume's sync wait stops at 30 seconds; Flash claims 45 ms.

Use asynchronous TTS whenever the audio is made ahead of the moment it is heard: ad voiceovers, narration for a video, audiobooks, announcements, per-page audio, and batches. Use a live, streaming voice only when a person is in a conversation and every 100 milliseconds is felt. Sume TTS 1.0 is the first kind: a job-based API with polling and signed webhooks, and no streaming output. Its sync mode is a bounded wait of at most 30 seconds, not a live stream.
Microsoft's own MAI-Voice-2.1 page makes the same split on its side. It lists the standard model at about 550 ms latency and recommends it for audiobooks, content creation and voice-overs, and lists the Flash model at about 45 ms for call centers, voice assistants and IVR. Suno's Speech beta, announced for everyone on 2026-10-01, sits on the made-ahead side too: you type, describe a voice and music, and get a finished track.
The four modes on a Sume TTS request
Every TTS request returns a job id right away. The mode field decides how you learn the outcome, and the docs recommend async or webhook for new integrations.
| Mode | What happens | Best for |
|---|---|---|
| async (default) | Returns at once with status, result and cancel URLs; you poll, honoring next_poll_after_seconds | Batches, scheduled jobs, any worker you control |
| webhook | Returns at once; Sume posts a signed job.completed, job.failed or job.canceled to your HTTPS URL | Servers that should not poll, many jobs in flight |
| sync | Blocks up to wait_timeout_seconds (0 to 30), then returns the job's current state | Short lines in a script, a cheap demo, a CLI |
| subscribe | An alias for sync: the same bounded wait, not a stream | Nothing new; use async or webhook instead |
Why 30 seconds is a ceiling, not a promise
A sync wait that expires is still a 2xx. The response carries the job's current state with polling URLs, and the docs say to keep polling rather than submit again, because a second request is a second paid job. Waiter capacity can also run out, in which case the response says so with a sync.capacity_exhausted marker. Both cases mean the same thing: the job exists and is still yours.
A long script is the case sync handles worst. A single job can synthesize up to 1,200 seconds of audio and the transcript limit is 20,000 characters, so a long narration job can outlast any wait. Submit with async, poll from your own process, and set the deadline in your client.
A short polling loop
The Node sketch below submits a job and polls until it is terminal, honoring the suggested delay. The envelope may be wrapped in data, so the code reads both forms. Reading result_ready is the clean stop condition, and a 409 from /result before that means the job is not completed yet.
const h = { Authorization: `Bearer ${process.env.SUME_API_KEY}`, "Content-Type": "application/json" };
const sub = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
method: "POST",
headers: { ...h, "Idempotency-Key": "async-demo-001" },
body: JSON.stringify({
transcript: "Your order has shipped.",
voice: { id: process.env.VOICE_ID },
language: "en",
mode: "async",
}),
});
const first = await sub.json();
const id = (first.data ?? first).request_id;
for (;;) {
const r = await fetch(`https://api.sume.com/v1/jobs/${id}/status`, { headers: h });
const s = await r.json();
const d = s.data ?? s;
if (d.terminal) { console.log(d.result_ready ? "ready" : "not ok"); break; }
await new Promise((ok) => setTimeout(ok, (d.next_poll_after_seconds ?? 2) * 1000));
}Choosing by job, not by taste
The cost side is the same in every mode: the price is per character sent, at $0.0475 per 1,000 on Sume as of 2026-10-07, and a retry with the same idempotency key is the same job. Mode changes only how you hear about the result.
- A voiceover for a video, an ad read or an audiobook chapter: async. Nobody is waiting, and you want idempotent retries.
- A thousand rows from a spreadsheet: webhook, or async with a bounded poller. Sume's queue accepts valid jobs and runs them within your workspace's concurrency limits, so
queuedis normal. - An IVR prompt bank or stream alert lines: async, ahead of time. Generate the lines once and play the files, which also removes latency from the call.
- A voice agent that must answer inside a conversation: not this API. Sume documents no streaming TTS, so use a live voice service for the call and Sume for the pre-rendered prompts around it.
- A button in an app that creates a short line on demand: sync is fine if the wait is short, but design the button to survive a timeout and fall back to polling.
Sources
Related posts
More in Developers
- Which Sume API calls are safe to retry blindly, and which need a key?
Reads, cancels and redelivers retry safely; paid submits retry only under the same Idempotency-Key. A call-by-call table, plus the codes that mean wait or stop.
- Which Sume job and run endings send a webhook, and which stay silent?
Jobs send job.completed, job.failed and job.canceled. Canceled or skipped runs send nothing. A matrix of terminal states and what your receiver can expect.
- Which Sume image models accept image_size for custom pixel dimensions?
Eleven of Sume's 19 image models accept image_size with a width and height: ChatGPT Image 2.5 and 2, Seedream, FLUX.2, Qwen and Recraft. Eight do not.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
Written by Sume