Edit a transcript with plain-language instructions: ElevenLabs STT

ElevenLabs STT edits a transcript from an instruction of up to 2,000 characters and returns edited_transcript. Sume captions align your own script_text.

4 min readSume
All posts

On Sept 28, 2026 the ElevenLabs changelog added transcript editing by natural-language instruction to its speech-to-text API: you describe the change in up to 2,000 characters and the response carries an edited_transcript. It is a convenient way to fix names, remove filler or tidy casing without writing find-and-replace rules. Sume has no equivalent operation in its docs. In Sume you supply the corrected text yourself and let captions align it to the audio.

What the changelog says

ElevenLabs changelog, read 2026-10-05
DateEntry
Sept 28STT transcript editing by natural-language instruction, instruction up to 2,000 characters, returns edited_transcript
Sept 28Eleven v4 and v4 Turbo; JS and Python SDK v2.70.0; CLI v1.4.0
Sept 11Scribe v2 Medical launched, same billing as Scribe v2

Using it safely

This post only confirms the edited_transcript response field, so read the ElevenLabs API reference before you build on it. Whatever the field names, the same cautions apply to any model-driven edit of a transcript:

  • Keep the original. Store the raw transcript next to the edited one and diff them, so a reviewer sees every change.
  • Keep instructions narrow: 'Spell the speaker name as in this list' is safer than 'make it better'. A broad instruction can change meaning.
  • Do not edit the words and then reuse the old timestamps. Timing belongs to the audio, so a changed word needs re-alignment.
  • Never edit a legal, medical or financial record by instruction without human sign-off.

The Sume route: your text, aligned

POST /v1/video-captions burns captions onto a public HTTPS video URL. Optional fields include style (slam, punch, tiktok-green, korean-ad), font, language (a speech-to-text hint), script_text, words, and cues/segments for authored overlay text. Per the docs, speech-to-captions works with script_text or, if you omit it, with STT, and needs audible speech: a silent clip fails with caption_no_speech.

So the Sume workflow is: get a transcript, correct it in your own editor or with your own model, send the corrected string as script_text, and let the captions follow the audio. The decision about what words are right stays with you, and the job is reproducible from the text you sent.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume