Clean a transcript with a 2,000-character instruction, then caption
ElevenLabs speech to text can edit a transcript from a natural-language instruction up to 2,000 characters. A cleanup then captions workflow with Sume captions.

Fix the transcript before you burn captions, not after. The ElevenLabs changelog for September 28, 2026, read 2026-10-03, says transcript editing takes a natural-language instruction of up to 2,000 characters; batch transcription returns an edited_transcript, and realtime emits an edited_transcript event. On Sume, the matching step is to pass corrected text as script_text to the captions job.
What the vendor added
The entry names three facts.
- The editing instruction is natural language, up to 2,000 characters.
- Batch results return
edited_transcript. - Realtime streams emit an
edited_transcriptevent.
The workflow
Generated video often carries speech with brand names, product codes and numbers that speech to text misreads. A cleanup pass is where you fix those, and a good instruction is specific: list the names with correct spellings, say whether to remove filler words, and say whether to keep numerals as digits. Keep it under 2,000 characters.
Captions on Sume
The video captions guide takes a public HTTPS video_url and returns a job-backed captioned video. Speech to captions uses script_text, or STT when you omit it, and needs audible speech: a silent clip fails as caption_no_speech with next_action: use_overlay_captions. If you already have authored lines, send cues or segments with text, start and end to burn overlay copy without speech recognition.
Omitted style follows the wording: slam for Latin copy and black-outline for Korean. A named style is rendered as named, and sending Korean copy to slam, punch or tiktok-green returns 400 rather than burning missing glyphs.
A starting instruction
A useful cleanup instruction is a short list: correct these names, spell these terms this way, remove filler words, keep numbers as digits, and do not paraphrase. Keep it factual and under the limit. Test it on one minute of audio and read the output before you run a batch.
Whatever tool you use, keep the original transcript next to the edited one so you can compare them.
Check what was heard first
Video inspect runs Sume STT 1.0 on a hosted clip when transcribe: true, at $0.01 per audio minute, and segmentation.mode: "sentence" also returns caption-shaped segments[]. Read that transcript, correct it, then send the corrected text as script_text. A silent clip fails with inspect_source_has_no_audio, so probe has_audio first.
Sources
Related posts
More in Use cases
- Article 50: creator duties vs provider duties, side by side
Article 50(4) puts deepfake and certain text disclosure on deployers, while 50(2) puts marking on providers. A two-column split for teams that publish AI video.
- Demand Gen carousels: 2 to 10 matching cards from one reference
Demand Gen carousels take 2 to 10 cards. Image assets run 4:5 or 9:16 at 5 MB. How to batch matching cards from one reference image with the Sume image API.
- Demand Gen logo 150 KB, images 5 MB: check sizes before upload
Demand Gen allows images up to 5 MB but logos only 150 KB. Check bytes on every Sume image before upload, since output compression is not served in v1.
- Demand Gen video ads: four ratios, a 5 s floor and the 4:5 gap
Demand Gen video takes 1:1, 16:9, 4:5 and 9:16, from 5 s. How to render the four-ratio set with Sume, including the 4:5 case the video API does not list.
Written by Sume