Check the transcript before captions: Scribe v2 error tiers
ElevenLabs lists Scribe v2 error rates by language group. Before burning captions on Sume, pass script_text or your own words so a wrong transcript never ships.

If you know what the speaker says, give the captions endpoint script_text so it aligns your text to the audio instead of trusting a transcript. ElevenLabs' own Scribe v2 documentation says accuracy varies a lot by language, so any pipeline that auto-transcribes should include a review step.
Auto-transcription is a good default. A caption you publish is a promise about what was said, so the default deserves a check.
What the vendor says about accuracy
The ElevenLabs speech-to-text doc says Scribe v2 covers 90+ languages. It groups them by word error rate: 34 languages at or under 5%, and 19 languages where the doc lists a rate above 25 to 50%. Files can reach 3 GB, and the doc lists keyterm prompting of up to 1000 terms.
Those are the vendor's published tiers. The doc does not give you a rate for your accent, microphone or topic, so a spot check on your own audio matters more than any table.
| Tier | Languages in the tier | Word error rate |
|---|---|---|
| Strong | 34 | 5% or lower |
| Weak | 19 | Above 25 to 50% |
Your options on Sume
The captions endpoint has a language field that is only a hint for speech recognition; it never picks the caption style or font. You have four routes, and they trade effort for control.
- Let it transcribe and review the output before publishing.
- Pass script_text to align a script you already have; script_alignment_mismatch or failed errors tell you when the audio and text disagree.
- Pass words, cues or segments to skip recognition entirely.
- Restyle an earlier caption with source_caption_id, which does not transcribe again.
A review habit that scales
Run the cheap path first, read the transcript for names, numbers and product terms, and only then render the final style. Keep a short list of brand terms and their spellings; the brand-name spelling post shows how script_text fixes them.
For clips that Sume generated itself, such as a TTS voiceover, you already have the source text and word timestamps, so recognition is unnecessary. Feeding those words straight into the captions request removes the transcription risk completely.
When the mismatch error is useful
A script_alignment_mismatch error is not a failure of your pipeline; it is a flag that the script and the recording differ. Fix the script, or re-record the section, and resubmit. Catching that before posting is cheaper than correcting a published caption.
Sources
Related posts
More in Media tools
- Children's captions at 17 characters per second: Netflix rule and Sume
Netflix's English style guide sets 20 chars per second for adults and 17 for children. See how to approximate it with Sume caption phrasing overrides.
- Cloudflare Images segment=foreground vs Sume RMBG PNG alpha
Cloudflare Images removes a background with segment=foreground, but its listed output formats lack PNG. Sume RMBG returns PNG artifacts with alpha.
- Cloudflare Stream animated GIF thumbnail vs a Sume silent MP4 clip
Cloudflare Stream builds a GIF from thumbnail.gif at 5 seconds and 8 fps by default. Sume makes no GIF; video trim returns a silent MP4 at 24 fps or more.
- Cloudflare Stream thumbnail fit modes vs Sume Timeline fit modes
Cloudflare Stream fit: crop, clip, scale, fill. Sume Timeline 1.0 fit: cover, contain, stretch, blur. Which names map and where they differ.
Written by Sume