Clean a transcript with a 2,000-character instruction, then caption

ElevenLabs speech to text can edit a transcript from a natural-language instruction up to 2,000 characters. A cleanup then captions workflow with Sume captions.

4 min readSume
All posts

Fix the transcript before you burn captions, not after. The ElevenLabs changelog for September 28, 2026, read 2026-10-03, says transcript editing takes a natural-language instruction of up to 2,000 characters; batch transcription returns an edited_transcript, and realtime emits an edited_transcript event. On Sume, the matching step is to pass corrected text as script_text to the captions job.

What the vendor added

The entry names three facts.

  • The editing instruction is natural language, up to 2,000 characters.
  • Batch results return edited_transcript.
  • Realtime streams emit an edited_transcript event.

The workflow

Generated video often carries speech with brand names, product codes and numbers that speech to text misreads. A cleanup pass is where you fix those, and a good instruction is specific: list the names with correct spellings, say whether to remove filler words, and say whether to keep numerals as digits. Keep it under 2,000 characters.

Captions on Sume

The video captions guide takes a public HTTPS video_url and returns a job-backed captioned video. Speech to captions uses script_text, or STT when you omit it, and needs audible speech: a silent clip fails as caption_no_speech with next_action: use_overlay_captions. If you already have authored lines, send cues or segments with text, start and end to burn overlay copy without speech recognition.

Omitted style follows the wording: slam for Latin copy and black-outline for Korean. A named style is rendered as named, and sending Korean copy to slam, punch or tiktok-green returns 400 rather than burning missing glyphs.

A starting instruction

A useful cleanup instruction is a short list: correct these names, spell these terms this way, remove filler words, keep numbers as digits, and do not paraphrase. Keep it factual and under the limit. Test it on one minute of audio and read the output before you run a batch.

Whatever tool you use, keep the original transcript next to the edited one so you can compare them.

Check what was heard first

Video inspect runs Sume STT 1.0 on a hosted clip when transcribe: true, at $0.01 per audio minute, and segmentation.mode: "sentence" also returns caption-shaped segments[]. Read that transcript, correct it, then send the corrected text as script_text. A silent clip fails with inspect_source_has_no_audio, so probe has_audio first.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume