ElevenLabs voice ids in the Sume TTS Router: list them, avoid 400s

Eleven v4 and other eleven-* ids run through the Sume TTS Router. How to list voices with family=eleven, which id shape each family takes, and what gets a 400.

4 min readSume
All posts

Yes: the Sume TTS Router can speak with ElevenLabs voices. Send model set to an eleven-* id and a voice.id that is an ElevenLabs voice id, and list your options first with GET /v1/tts-router/voices?family=eleven.

Older write-ups, and the OpenAPI snapshot published with the docs at the time of writing, show only Sonic ids. The router code lists twelve ids, so call GET /v1/tts-router/models to see what your deployment serves before you hard-code anything.

The eleven ids and their request limits

Every router request needs a model. The ElevenLabs rows below are the ones in the router catalog code. Billing is the list price times 1.25, rounded up to the next cent per request, so the per-1,000-character column is the price before that rounding.

Eleven rows of the TTS Router catalog, from the Sume catalog code, read 2026-10-11
Model idBilled per 1,000 charactersMax characters per requestLanguages
eleven-v4$0.1010,00044
eleven-v4-turbo$0.0510,00044
eleven-v3$0.1255,00044
eleven-multilingual-v2$0.12510,00028
eleven-turbo-v2.5$0.062520,00031

List the voices before you send text

GET /v1/tts-router/voices requires a family query parameter, either sonic or eleven. The answer is cached for ten minutes and wrapped in data, with family, a source field and a voices array. source is live when the voice list was read from an ElevenLabs account, and premade when Sume falls back to its built-in set of 21 premade voices.

Each voice carries id, name, description, language, gender, accent, preview_url and tags. Use id as voice.id in the generate call. If the provider is not configured or cannot be reached, the endpoint answers 503 with provider_not_configured or provider_unavailable; retry later rather than guessing an id.

import os, json, urllib.request

BASE = "https://api.sume.com"
HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

req = urllib.request.Request(
    f"{BASE}/v1/tts-router/voices?family=eleven", headers=HEAD
)
with urllib.request.urlopen(req) as resp:
    data = json.load(resp)["data"]

print("source:", data["source"])
for voice in data["voices"][:10]:
    print(voice["id"], voice["name"], voice["accent"])

Which voice id shape each family accepts

The Sonic and Eleven families do not share an id format, and mixing them is the commonest cause of a 400.

  • Sonic rows take a TTS voice UUID, or a voi_ followed by 32 hex characters for a voice from your Voices library. Any other shape returns 400 invalid_voice_id.
  • Eleven rows take an ElevenLabs voice id, a 20-character string of letters and digits, taken from the family=eleven list.
  • Sonic's list shows only public voices. Your own clones stay in the Voices library and are reached through voi_* ids on sonic-* models.
  • sonic-preview is not compatible with pro voice clones and returns voice_model_mismatch.

What an eleven row refuses

The Eleven rows accept voice_settings, apply_text_normalization and, on v4 and v4-turbo, a seed. They return 400 for fields that only exist on the Sonic path: avatar_id, avatar_handle, transcript_source, generation_config, speed, timestamps, segmentation and pronunciation_dict_id. Output is mp3 only, in the 44.1 kHz family, with mp3_44100_128 as the default.

Two consequences follow. A script that asks Sonic for word timestamps cannot simply swap the model string, and a pronunciation dictionary built for Sonic does not apply. Plan those two features separately when you move a Sonic workflow to an Eleven voice.

Voice cloning remains an app feature. The API does not create clones, so the ids you list here are the ones already available to you.

A safe way to adopt a new id

Read GET /v1/tts-router/models first, pick an id and its character limit, list voices for the matching family, and only then send text. The API reference covers authentication, and the models overview links the audio endpoints. Because pricing is rounded up per request, batching short lines into fewer requests keeps the rounding from dominating the bill.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume