Check the transcript before captions: Scribe v2 error tiers

ElevenLabs lists Scribe v2 error rates by language group. Before burning captions on Sume, pass script_text or your own words so a wrong transcript never ships.

4 min readSume
All posts

If you know what the speaker says, give the captions endpoint script_text so it aligns your text to the audio instead of trusting a transcript. ElevenLabs' own Scribe v2 documentation says accuracy varies a lot by language, so any pipeline that auto-transcribes should include a review step.

Auto-transcription is a good default. A caption you publish is a promise about what was said, so the default deserves a check.

What the vendor says about accuracy

The ElevenLabs speech-to-text doc says Scribe v2 covers 90+ languages. It groups them by word error rate: 34 languages at or under 5%, and 19 languages where the doc lists a rate above 25 to 50%. Files can reach 3 GB, and the doc lists keyterm prompting of up to 1000 terms.

Those are the vendor's published tiers. The doc does not give you a rate for your accent, microphone or topic, so a spot check on your own audio matters more than any table.

Scribe v2 accuracy tiers as documented (read 2026-10-03)
TierLanguages in the tierWord error rate
Strong345% or lower
Weak19Above 25 to 50%

Your options on Sume

The captions endpoint has a language field that is only a hint for speech recognition; it never picks the caption style or font. You have four routes, and they trade effort for control.

  • Let it transcribe and review the output before publishing.
  • Pass script_text to align a script you already have; script_alignment_mismatch or failed errors tell you when the audio and text disagree.
  • Pass words, cues or segments to skip recognition entirely.
  • Restyle an earlier caption with source_caption_id, which does not transcribe again.

A review habit that scales

Run the cheap path first, read the transcript for names, numbers and product terms, and only then render the final style. Keep a short list of brand terms and their spellings; the brand-name spelling post shows how script_text fixes them.

For clips that Sume generated itself, such as a TTS voiceover, you already have the source text and word timestamps, so recognition is unnecessary. Feeding those words straight into the captions request removes the transcription risk completely.

When the mismatch error is useful

A script_alignment_mismatch error is not a failure of your pipeline; it is a flag that the script and the recording differ. Fix the script, or re-record the section, and resubmit. Catching that before posting is cheaper than correcting a published caption.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume