Port an ElevenLabs text-to-speech call to Sume TTS 1.0, field by field

ElevenLabs puts the voice in the URL; Sume TTS 1.0 takes it in the body and rejects model_id. A field map and a tested mapper function.

6 min readSume
All posts

Moving a text-to-speech call from ElevenLabs to Sume TTS 1.0 changes five things: the auth header, where the voice goes, the audio format field, the removal of model_id, and the response, which becomes a job you poll instead of audio bytes. The voice is the biggest change, because an ElevenLabs voice id does not exist on Sume and must be replaced by a Sume avatar or voice.

ElevenLabs details are from its Create speech page, read 2026-10-10. Sume details are from the OpenAPI description of POST /v1/tts-1.0/generate and the Jobs and results guide.

The two requests side by side

ElevenLabs sends POST https://api.elevenlabs.io/v1/text-to-speech/{voice_id} with an xi-api-key header, text in the body, a default model of eleven_multilingual_v2, and an output_format string such as mp3_44100_128, opus_48000_64 or wav_48000. Sume sends POST https://api.sume.com/v1/tts-1.0/generate with Authorization: Bearer or x-api-key (never both), a transcript of up to 20,000 characters, and a voice chosen in the body.

ElevenLabs fields from its Create speech page (read 2026-10-10); Sume fields from the TTS 1.0 OpenAPI description.
ConcernElevenLabsSume TTS 1.0
Auth headerxi-api-keyAuthorization: Bearer or x-api-key, exactly one
Voicevoice_id in the pathavatar_handle or avatar_id at the top level, or voice.id (a Sume UUID or voi_ id)
TextText field in the bodytranscript, 1 to 20,000 characters
Model choicemodel_id, default eleven_multilingual_v2None. model and model_id are rejected with 400
FormatA string like mp3_44100_128An object: container, sample_rate, bit_rate, encoding
Default formatNot read from the pagemp3, 44100 Hz, 128000 bit/s
ResultAudio in the responseA job: poll status_url, then read result_url
Length ceilingNot read from the pageAudio over 1200 s fails with tts_duration_exceeded

A mapper you can run

This function turns an ElevenLabs-style body and format string into a Sume body. It refuses formats Sume's schema has no container for, rather than guess, and it needs a Sume avatar handle because the old voice id cannot carry over.

import re

def to_sume_tts(body, output_format, avatar_handle):
    m = re.fullmatch(r"(mp3|wav|pcm)_(\d+)(?:_(\d+))?", output_format)
    if not m:
        raise ValueError(f"no Sume container for {output_format}")
    kind, rate, kbps = m.group(1), int(m.group(2)), m.group(3)
    if rate not in (8000, 16000, 22050, 24000, 44100, 48000):
        raise ValueError(f"unsupported sample rate {rate}")
    fmt = {"container": "raw" if kind == "pcm" else kind, "sample_rate": rate}
    if kind == "mp3":
        fmt["bit_rate"] = int(kbps or 128) * 1000
    out = {"transcript": body["text"], "avatar_handle": avatar_handle,
           "output_format": fmt}
    return out

print(to_sume_tts({"text": "Hello there."}, "mp3_44100_128", "@narrator"))
try:
    to_sume_tts({"text": "x"}, "opus_48000_64", "@narrator")
except ValueError as e:
    print(e)

What does not carry over

Sume's TTS 1.0 has no engine picker. If you must choose a specific catalog model, the docs point to the separate TTS Router at POST /v1/tts-router/generate. Voice ids are the other break: a name or id from another ecosystem is rejected with 400 before a job is queued. Pick a ready avatar from GET /v1/avatar-1.0/avatars, where voice.status is ready, and use its handle. Voice cloning is not in the public API.

Send an Idempotency-Key on the create and retry with the same key after a timeout, so a repeat does not bill twice. For word timings, ask for timestamps.words; for per-sentence audio slices, add segmentation.mode: sentence.

Where ElevenLabs is the better fit

If you depend on a particular ElevenLabs voice, on its model choices, or on audio streaming back in the same HTTP response, the move costs you more than the format fields. Sume's route is job-shaped: phase one is async job plus poll or webhook, non-streaming. Move when you want TTS to sit in one account with video, timeline and avatar jobs, and test a few of your real scripts before you cut over.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume