Edit a transcript with plain-language instructions: ElevenLabs STT
ElevenLabs STT edits a transcript from an instruction of up to 2,000 characters and returns edited_transcript. Sume captions align your own script_text.

On Sept 28, 2026 the ElevenLabs changelog added transcript editing by natural-language instruction to its speech-to-text API: you describe the change in up to 2,000 characters and the response carries an edited_transcript. It is a convenient way to fix names, remove filler or tidy casing without writing find-and-replace rules. Sume has no equivalent operation in its docs. In Sume you supply the corrected text yourself and let captions align it to the audio.
What the changelog says
| Date | Entry |
|---|---|
| Sept 28 | STT transcript editing by natural-language instruction, instruction up to 2,000 characters, returns edited_transcript |
| Sept 28 | Eleven v4 and v4 Turbo; JS and Python SDK v2.70.0; CLI v1.4.0 |
| Sept 11 | Scribe v2 Medical launched, same billing as Scribe v2 |
Using it safely
This post only confirms the edited_transcript response field, so read the ElevenLabs API reference before you build on it. Whatever the field names, the same cautions apply to any model-driven edit of a transcript:
- Keep the original. Store the raw transcript next to the edited one and diff them, so a reviewer sees every change.
- Keep instructions narrow: 'Spell the speaker name as in this list' is safer than 'make it better'. A broad instruction can change meaning.
- Do not edit the words and then reuse the old timestamps. Timing belongs to the audio, so a changed word needs re-alignment.
- Never edit a legal, medical or financial record by instruction without human sign-off.
The Sume route: your text, aligned
POST /v1/video-captions burns captions onto a public HTTPS video URL. Optional fields include style (slam, punch, tiktok-green, korean-ad), font, language (a speech-to-text hint), script_text, words, and cues/segments for authored overlay text. Per the docs, speech-to-captions works with script_text or, if you omit it, with STT, and needs audible speech: a silent clip fails with caption_no_speech.
So the Sume workflow is: get a transcript, correct it in your own editor or with your own model, send the corrected string as script_text, and let the captions follow the audio. The decision about what words are right stays with you, and the job is reproducible from the text you sent.
Sources
Related posts
More in Developers
- Embargoed Black Friday reveal: Sume media URLs are public, copy first
Format artifacts sit on durable public media.sume.com URLs with no expiry. For an embargoed drop, copy the file to your own storage and never log the URL.
- Emotion tags in the script or an emotion field: Sume's bill
MAI-Voice takes emotion tags. Sume's documented control is generation_config.emotion, and every transcript character is billed, tags included.
- 'end_image_url requires image_url': a last frame needs a first frame
A last frame without a first frame is refused on both Sume video routes. Add a first frame, or drop the end frame and describe the ending in the prompt.
- An es-MX voice with an es-ES request: Sume compares primary language
Sume's TTS language check compares regional tags by primary language, so es-MX against es-ES passes. What that means for accent, and what 409 you get otherwise.
Written by Sume