ElevenLabs product list vs Sume API routes: what has no equivalent
Sume's OpenAPI has speech, transcription, music and word timings, but no dubbing, voice isolation, sound effects, voice changer or voice cloning route.

Sume covers text to speech, speech to text, music and word timings in its API, but it has no route for dubbing, voice isolation, sound effects, a voice changer or voice cloning. If your ElevenLabs integration depends on any of those, Sume does not replace it today.
The ElevenLabs list is from its documentation overview, read 2026-10-10, which names text to speech, speech to text, music, text to dialogue, image and video, voice changer, voice isolator, dubbing, sound effects, voice cloning and design, remixing, forced alignment and agents. The Sume side is from the OpenAPI document published with the docs, where I listed every path, plus the Timeline audio and Music Router pages.
Matching the list to Sume's routes
I checked each ElevenLabs capability against the paths in Sume's OpenAPI file. A match means a route exists for the same kind of work, not that the output is equal. Quality, languages and voices are things to test yourself.
| ElevenLabs capability | Sume route found | Note |
|---|---|---|
| Text to speech | POST /v1/tts-1.0/generate, /v1/tts-router/generate | Async job; router picks a catalog model |
| Speech to text | POST /v1/stt-1.0/transcribe | Word timings always returned |
| Music | /v1/music-1.0/generate, /v1/music-router/generate | Also a BGM catalog |
| Forced alignment | Partly: TTS timestamps.words | Timings come from text Sume itself synthesizes |
| Image and video | /v1/images, /v1/videos | Separate model catalogs |
| Agents | /v1/agent/completions, /v1/agent-runs | Different product shape |
| Dubbing | None found | |
| Voice isolator | None found | /v1/audio-detach separates audio from video, which is not noise removal |
| Sound effects | None found | |
| Voice changer | None found | |
| Voice cloning and design | Not in the public API | Cloning is an app feature only |
Close calls
Two rows look like matches and are not. /v1/audio-detach pulls the audio track out of a video file; it does not clean speech. The TTS timestamps are word timings for audio Sume generates, so they are not an alignment service for arbitrary audio you bring. For arbitrary audio, STT 1.0 gives you words[] with start and end times, which covers many alignment needs, but it is transcription, not alignment of a known script.
Dubbing is the largest gap. You can assemble a rough equivalent from STT, a translation step you supply, TTS and Timeline, but that is a pipeline you maintain, with no lip-sync or voice matching from Sume. An ElevenLabs dubbing call is one request.
Where Sume fits instead
The reason to use Sume for audio is proximity to the rest of a video job. The same key and the same job envelope cover TTS, music, avatar video, Timeline renders and captions, and a Timeline can take a TTS result as its audio spine. If your product only needs a voice or a sound-effects library, a specialist audio vendor will serve you better, and this post is not a reason to move.
Before you decide, list the ElevenLabs calls your code makes today and mark each against the table. If all of them have a route on the Sume side, run a small test of your real scripts. If one has none, keep that one where it is and move only the rest.
Sources
Related posts
More in Comparisons
- Scribe v2 audio-event tags and 32 speakers vs Sume STT's fixed flags
ElevenLabs Scribe v2 tags audio events and separates up to 32 speakers. Sume STT fixes diarize and tag_audio_events server-side, so those toggles do not exist.
- ElevenLabs sound effects: prompt influence, 48 kHz WAV, and Sume
ElevenLabs SFX offers high or low prompt influence and 48 kHz WAV. Sume has no sound-effects route; a Music prompt plus a split makes a sting. What you lose.
- Gemini TTS has 30 prebuilt voices; Sume TTS takes avatar voice ids
Google's Gemini TTS page lists 30 prebuilt voices. Sume TTS takes a voice UUID or voi_ id, or an avatar whose voice.status is ready; cloning stays app-only.
- Grok Imagine's 4 keyframes and 7 references vs Sume's one image
xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.
Written by Sume