ElevenLabs API timeout: cascade_timeout_seconds vs Sume waits
ElevenLabs added cascade_timeout_seconds (2-15 s) for Speech Engine retries. Sume's wait_timeout_seconds is a different knob: a 0-30 s HTTP wait on a job.

The newest ElevenLabs timeout setting is cascade_timeout_seconds, a number from 2 to 15 (default 4) on Speech Engine create and update requests. It sets how long ElevenLabs waits for an upstream speech engine before retrying. It is not a limit on your own HTTP call. Sume's nearest setting, wait_timeout_seconds, is a different thing: how long a submit call blocks for a job, clamped to 0..30.
ElevenLabs details are from its 2026-09-21 changelog entry; Sume details are from Jobs and results, both read 2026-10-01.
What does cascade_timeout_seconds control?
The changelog describes it as the control for how long ElevenLabs waits for an upstream speech engine before retrying. The accepted range is 2 to 15 seconds and the default is 4. It appears on Speech Engine create and update requests, so it is a property you configure on the engine, not a client-side deadline for a single call. The entry says nothing about request or read timeouts on other ElevenLabs endpoints, so this post does not either.
What does wait_timeout_seconds control on Sume?
Every Sume submit endpoint takes a mode. With sync, the server holds the HTTP request for up to wait_timeout_seconds waiting for a terminal state. The docs state the value is clamped to 0..30 and bounds how long the request blocks, not how long the job may take.
If the budget runs out, the response is still a 2xx with the job id, and you continue with the status URL. The docs' instruction for that case is "Poll. Do not resubmit."
| ElevenLabs `cascade_timeout_seconds` | Sume `wait_timeout_seconds` | |
|---|---|---|
| Range | 2 to 15, default 4 | Clamped to 0..30 |
| What it bounds | Wait for an upstream speech engine before a retry | How long the submit HTTP request blocks |
| Set on | Speech Engine create and update | Any submit call in sync mode |
| When it expires | ElevenLabs retries | You still get the job id; poll the status URL |
What happens past 30 seconds on Sume?
Nothing is lost. The job keeps running and the envelope carries status_url, result_url and a sync object whose timed_out flag tells you the wait ended early. Poll GET /v1/jobs/:id/status with backoff until terminal is true, then read the result. Resubmitting the create would pay for a second job; if you must retry the submit itself, reuse the same Idempotency-Key.
Which one should my client set?
Treat them as separate layers. If you use ElevenLabs Speech Engine, cascade_timeout_seconds is an engine setting to tune. If you call Sume, prefer mode: "async" and poll, so your own client timeout can be as long as the work needs. See Sume API timeouts for the client side, and subscribe is a sync alias for why a longer wait is not available.
Sources
Related posts
More in Developers
- ElevenLabs API cursor pagination, and how Sume pages lists
ElevenLabs added a cursor-paginated phone number endpoint. Sume list routes use keyset pages: pass next_cursor back as cursor until has_more is false.
- ElevenLabs Flows webhook: what changed, and Sume's per-job webhook
ElevenLabs Flows template runs can send a terminal result to a webhook, and webhook targets now need type all or ids. Sume sets the callback per job.
- is_final_audio_for_turn vs Sume TTS sentence boundaries: wav or mp3
ElevenLabs now emits is_final_audio_for_turn after every byte, even for MP3. Sume marks boundaries differently: sentence segments, sliced only for wav or raw.
- eleven_turbo_v2_5 deprecated: use Flash; Sume TTS Router ids
ElevenLabs lists eleven_turbo_v2_5 as deprecated and suggests eleven_flash_v2_5. Sume's TTS Router pins an explicit catalog id, so list the ids first.
Written by Sume