ElevenLabs speech to speech API: what Sume offers instead
ElevenLabs lists speech-to-speech voice changer models. Sume has no voice-to-voice endpoint: transcribe with STT, edit the text, then run TTS.

Sume has no speech-to-speech endpoint. Its text-to-speech request takes a transcript, not source audio, so the nearest route is three steps: transcribe with STT 1.0, edit the text, then synthesize it with TTS 1.0. The new voice follows the words, not the original performance.
The ElevenLabs models page (read 2026-10-01) lists eleven_multilingual_sts_v2 as a multilingual voice changer (Speech to Speech) and eleven_english_sts_v2 as an English-only one. Sume facts are from its OpenAPI file and Timeline audio docs.
What does a speech-to-speech model do?
On the vendor's page these are voice changer models: audio in, audio out in a different voice. Sume does not list an equivalent, so nothing in the Sume API takes a recorded take and re-voices it.
What does Sume TTS take as input?
The TTS 1.0 schema says "Exactly one of transcript_source or transcript is required." Text goes in; audio comes out. The voice is picked by the top-level avatar_id / avatar_handle, or voice.id.
STT 1.0 is the route that supplies the words: the schema describes POST /v1/stt-1.0/transcribe as the primary Sume STT 1.0 route for speech-to-text.
| Step | Sume route | What carries over |
|---|---|---|
| Transcribe the take | POST /v1/stt-1.0/transcribe | Words and word timings |
| Edit the text | Your code or an editor | Wording only |
| Synthesize | POST /v1/tts-1.0/generate | A new voice and new delivery |
What is lost compared with voice conversion?
Pacing, emphasis and breath from the original take are not carried into TTS; only the transcript is. If the original delivery matters, keep it and do not use this route.
How do I put the new audio back on a video?
Timeline audio returns one hosted file, and the docs say to use that file as audio.url on the render. For lip-sync limits when the picture is a talking face, read replace video audio with an AI voice. For the full transcribe, translate, speak chain see build an AI dubbing pipeline.
Sources
Related posts
More in Models
- FLUX.2 flex steps and guidance: BFL has them, Sume does not
BFL lists adjustable steps and guidance only for FLUX.2 flex ($0.06/MP). Sume lists flux.2-flex but rejects unlisted parameters with 400.
- FLUX.2 max grounding search: what it is, what Sume lists
Only FLUX.2 [max] does web-grounded generation at BFL. Sume lists flux.2-pro and flux.2-flex, so grounding is not available there.
- GPT Image aspect ratios for a Pinterest pin (2:3) and Instagram (4:5)
Sume lists 17 aspect ratios, including 2:3 for a Pinterest pin and 4:5 for an Instagram portrait. A model only accepts the ratios its catalog row lists.
- HappyHorse 1.0 API: what Runway lists and how to check Sume's models
Runway lists HappyHorse 1.0 at 3-15 seconds, 720p and 1080p. Sume's video docs name no such id, so read /v1/videos/models before you plan around it.
Written by Sume