ElevenLabs speech to speech API: what Sume offers instead

ElevenLabs lists speech-to-speech voice changer models. Sume has no voice-to-voice endpoint: transcribe with STT, edit the text, then run TTS.

4 min readSume
All posts

Sume has no speech-to-speech endpoint. Its text-to-speech request takes a transcript, not source audio, so the nearest route is three steps: transcribe with STT 1.0, edit the text, then synthesize it with TTS 1.0. The new voice follows the words, not the original performance.

The ElevenLabs models page (read 2026-10-01) lists eleven_multilingual_sts_v2 as a multilingual voice changer (Speech to Speech) and eleven_english_sts_v2 as an English-only one. Sume facts are from its OpenAPI file and Timeline audio docs.

What does a speech-to-speech model do?

On the vendor's page these are voice changer models: audio in, audio out in a different voice. Sume does not list an equivalent, so nothing in the Sume API takes a recorded take and re-voices it.

What does Sume TTS take as input?

The TTS 1.0 schema says "Exactly one of transcript_source or transcript is required." Text goes in; audio comes out. The voice is picked by the top-level avatar_id / avatar_handle, or voice.id.

STT 1.0 is the route that supplies the words: the schema describes POST /v1/stt-1.0/transcribe as the primary Sume STT 1.0 route for speech-to-text.

Steps for the workaround, checked 2026-10-01.
StepSume routeWhat carries over
Transcribe the takePOST /v1/stt-1.0/transcribeWords and word timings
Edit the textYour code or an editorWording only
SynthesizePOST /v1/tts-1.0/generateA new voice and new delivery

What is lost compared with voice conversion?

Pacing, emphasis and breath from the original take are not carried into TTS; only the transcript is. If the original delivery matters, keep it and do not use this route.

How do I put the new audio back on a video?

Timeline audio returns one hosted file, and the docs say to use that file as audio.url on the render. For lip-sync limits when the picture is a talking face, read replace video audio with an AI voice. For the full transcribe, translate, speak chain see build an AI dubbing pipeline.

Sources

Related posts

More in Models

All Models posts

Written by Sume