ElevenLabs product list vs Sume API routes: what has no equivalent

Sume's OpenAPI has speech, transcription, music and word timings, but no dubbing, voice isolation, sound effects, voice changer or voice cloning route.

6 min readSume
All posts

Sume covers text to speech, speech to text, music and word timings in its API, but it has no route for dubbing, voice isolation, sound effects, a voice changer or voice cloning. If your ElevenLabs integration depends on any of those, Sume does not replace it today.

The ElevenLabs list is from its documentation overview, read 2026-10-10, which names text to speech, speech to text, music, text to dialogue, image and video, voice changer, voice isolator, dubbing, sound effects, voice cloning and design, remixing, forced alignment and agents. The Sume side is from the OpenAPI document published with the docs, where I listed every path, plus the Timeline audio and Music Router pages.

Matching the list to Sume's routes

I checked each ElevenLabs capability against the paths in Sume's OpenAPI file. A match means a route exists for the same kind of work, not that the output is equal. Quality, languages and voices are things to test yourself.

ElevenLabs capabilities from its overview page (read 2026-10-10); Sume routes from the published OpenAPI paths.
ElevenLabs capabilitySume route foundNote
Text to speechPOST /v1/tts-1.0/generate, /v1/tts-router/generateAsync job; router picks a catalog model
Speech to textPOST /v1/stt-1.0/transcribeWord timings always returned
Music/v1/music-1.0/generate, /v1/music-router/generateAlso a BGM catalog
Forced alignmentPartly: TTS timestamps.wordsTimings come from text Sume itself synthesizes
Image and video/v1/images, /v1/videosSeparate model catalogs
Agents/v1/agent/completions, /v1/agent-runsDifferent product shape
DubbingNone found
Voice isolatorNone found/v1/audio-detach separates audio from video, which is not noise removal
Sound effectsNone found
Voice changerNone found
Voice cloning and designNot in the public APICloning is an app feature only

Close calls

Two rows look like matches and are not. /v1/audio-detach pulls the audio track out of a video file; it does not clean speech. The TTS timestamps are word timings for audio Sume generates, so they are not an alignment service for arbitrary audio you bring. For arbitrary audio, STT 1.0 gives you words[] with start and end times, which covers many alignment needs, but it is transcription, not alignment of a known script.

Dubbing is the largest gap. You can assemble a rough equivalent from STT, a translation step you supply, TTS and Timeline, but that is a pipeline you maintain, with no lip-sync or voice matching from Sume. An ElevenLabs dubbing call is one request.

Where Sume fits instead

The reason to use Sume for audio is proximity to the rest of a video job. The same key and the same job envelope cover TTS, music, avatar video, Timeline renders and captions, and a Timeline can take a TTS result as its audio spine. If your product only needs a voice or a sound-effects library, a specialist audio vendor will serve you better, and this post is not a reason to move.

Before you decide, list the ElevenLabs calls your code makes today and mark each against the table. If all of them have a route on the Sume side, run a small test of your real scripts. If one has none, keep that one where it is and move only the rest.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume