ElevenLabs API alternative: what to check before switching TTS

Thinking of moving text to speech off ElevenLabs? Sume's TTS is an async job API with Sonic ids on a router. The contract points to check before you switch.

4 min readSume
All posts

Sume can serve as an ElevenLabs API alternative for text to speech only if its contract fits your code: it is an asynchronous job API, its router lists five Sonic model ids and no Eleven id, and voices are Sume voice ids rather than voice names from elsewhere. Read the seven points below against your integration before you move traffic.

What is different in Sume's request contract?

Everything in this table comes from Sume's own API reference. ElevenLabs says its Eleven v4 models are available through its own API; nothing here compares quality or price with it.

Points to check when moving TTS calls to Sume, read 2026-09-29.
PointSume TTS contract
Authx-api-key or Authorization: Bearer
DeliveryAsync job with poll or webhook; sync waits at most 30 seconds
Lengthtranscript up to 20,000 characters
VoiceSume TTS voice UUID or voi_ library id, or an avatar's voice
EngineTTS 1.0 has no picker; the router takes sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview
Outputmp3 by default at 44,100 Hz and 128,000 bit rate; wav and raw available
BillingPer character; router list price times 1.10

Does the delivery model change my code?

Yes, if you consume audio as a stream. The reference describes Phase 1 as an async job with poll or webhook and non-streaming. A mode: "sync" request only blocks for up to wait_timeout_seconds, capped at 30, then returns polling URLs; the job keeps running. Plan for a job id, a status URL and a result URL that returns audio artifacts on media.sume.com.

What happens to my existing voices and tags?

A voice id that is not a Sume UUID or voi_ id is rejected with 400 before any job is queued or credit reserved, so existing voice ids need mapping to Sume voices first. Inline delivery tags in the text are also not part of the schema; delivery is set through generation_config with speed from 0.6 to 1.5, volume from 0.5 to 2 and an emotion string.

What does Sume not tell me?

Sume's schema does not publish a language count, a latency figure or a voice-cloning flow for the router, and this post states none. Latency claims such as the ~150 ms time to first speech ElevenLabs gives for Eleven v4 Turbo describe that vendor's own service; they do not carry over to a non-streaming job API. If your product needs streamed audio the moment text arrives, an async job API is a poor fit, and the honest answer is to keep that path where it is.

If you generate narration, voiceovers or dubbing tracks that are stitched into video afterwards, the job model matters less, because you wait for the file anyway. Sume's Timeline audio and video routes then take the finished audio as a durable URL. See TTS API for video narration for that shape.

How do I try one call?

Send one short transcript through the router and read the result URL when the job completes.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Your order has shipped.",
    "avatar_handle": "product_host",
    "mode": "async"
  }'

Sources

Related posts

More in Developers

All Developers posts

Written by Sume