A 2.5% word error rate is 25 wrong words per 1,000: plan the proofread

If a 2.5% WER held on your audio, a 1,000-word transcript would still carry about 25 wrong words. How to proofread and burn captions on Sume with script_text.

5 min readSume
All posts

The arithmetic

A 2.5 percent word error rate sounds close to perfect until you apply it to a transcript you will publish: 25 wrong words in every 1,000. That is the arithmetic for a rate you do not control, and it is the reason captions you burn onto video should go through a review step even when the transcription model is at the top of a leaderboard.

The 2.5 percent figure is the final-transcript rate for MAI-Transcribe-2-Streaming on the Artificial Analysis streaming leaderboard dated September 28, 2026, as relayed by Unite.AI (read 2026-10-07). The same chart lists 2.8 percent for the first partial. Your own audio will score differently, so use the figures below as a planning example, not as a promise.

Word count is the unit, so convert your clip length with a speaking-rate you measure yourself from a sample, not one you assume. The table uses word counts directly.

Wrong words expected if a 2.5% rate held (arithmetic on the reported rate, read 2026-10-07)
Transcript lengthWrong words at 2.5%Wrong words at 2.8% (first-partial rate)
250 words6.257
1,000 words2528
2,500 words62.570

Which errors matter

Not every error costs the same. A wrong "the" is invisible in a caption. A wrong product name, price or person's name is the one that gets screenshotted. Proofread the nouns first.

When you have the script: script_text

If you already have the script, do not proofread the transcript at all. Send your text as script_text on the standalone captions job. Per the video captions docs, Sume then keeps the speech-to-text word timings as the source of truth for time and aligns your text to them, so the wording is yours and the timing is the speech's.

Alignment can fail, and the failure is typed. The docs name script_alignment_mismatch and script_alignment_failed, with the recommended next action simplify_script_text_or_omit. If you omit script_text, Sume burns the speech-to-text text. The alignment mismatch fix covers the usual causes.

When you do not: fix once, burn many

When you do not have a script, correct the transcript once and reuse it. The captions job accepts words, word-level text, start and end in seconds, and in that form Sume does not run speech-to-text again. You can send only one of script_text, words, cues and segments. To restyle later, send source_caption_id instead of video_url, which reuses the word timings Sume already has.

What it costs

A standalone caption job reserves and captures $0.20 for a video of up to 60 seconds, per the same docs, and the live price is on GET /v1/catalog. Transcribing the same audio on its own costs $0.01 per audio minute at the public rate. The proofreading time is the expensive part, so spend it where it counts: names, numbers and anything a legal reviewer would read aloud.

Planning the proofread

A rate is only useful once it turns into an hour of work. At 25 wrong words per 1,000, a 5,000-word transcript has about 125 errors, and a reviewer reading at a normal pace will need to listen as well as read, because errors in names and numbers look plausible on the page.

Put the review where mistakes cost most: names, figures, dates and quoted statements. Sample a few minutes first and count errors; if the rate on your audio is higher than the headline number, plan more time before you start.

  • Names, numbers and dates first.
  • Sample a few minutes to get your own rate.
  • Listen, do not only read.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume