A 2.5% word error rate is 25 wrong words per 1,000: plan the proofread
If a 2.5% WER held on your audio, a 1,000-word transcript would still carry about 25 wrong words. How to proofread and burn captions on Sume with script_text.

The arithmetic
A 2.5 percent word error rate sounds close to perfect until you apply it to a transcript you will publish: 25 wrong words in every 1,000. That is the arithmetic for a rate you do not control, and it is the reason captions you burn onto video should go through a review step even when the transcription model is at the top of a leaderboard.
The 2.5 percent figure is the final-transcript rate for MAI-Transcribe-2-Streaming on the Artificial Analysis streaming leaderboard dated September 28, 2026, as relayed by Unite.AI (read 2026-10-07). The same chart lists 2.8 percent for the first partial. Your own audio will score differently, so use the figures below as a planning example, not as a promise.
Word count is the unit, so convert your clip length with a speaking-rate you measure yourself from a sample, not one you assume. The table uses word counts directly.
| Transcript length | Wrong words at 2.5% | Wrong words at 2.8% (first-partial rate) |
|---|---|---|
| 250 words | 6.25 | 7 |
| 1,000 words | 25 | 28 |
| 2,500 words | 62.5 | 70 |
Which errors matter
Not every error costs the same. A wrong "the" is invisible in a caption. A wrong product name, price or person's name is the one that gets screenshotted. Proofread the nouns first.
When you have the script: script_text
If you already have the script, do not proofread the transcript at all. Send your text as script_text on the standalone captions job. Per the video captions docs, Sume then keeps the speech-to-text word timings as the source of truth for time and aligns your text to them, so the wording is yours and the timing is the speech's.
Alignment can fail, and the failure is typed. The docs name script_alignment_mismatch and script_alignment_failed, with the recommended next action simplify_script_text_or_omit. If you omit script_text, Sume burns the speech-to-text text. The alignment mismatch fix covers the usual causes.
When you do not: fix once, burn many
When you do not have a script, correct the transcript once and reuse it. The captions job accepts words, word-level text, start and end in seconds, and in that form Sume does not run speech-to-text again. You can send only one of script_text, words, cues and segments. To restyle later, send source_caption_id instead of video_url, which reuses the word timings Sume already has.
What it costs
A standalone caption job reserves and captures $0.20 for a video of up to 60 seconds, per the same docs, and the live price is on GET /v1/catalog. Transcribing the same audio on its own costs $0.01 per audio minute at the public rate. The proofreading time is the expensive part, so spend it where it counts: names, numbers and anything a legal reviewer would read aloud.
Planning the proofread
A rate is only useful once it turns into an hour of work. At 25 wrong words per 1,000, a 5,000-word transcript has about 125 errors, and a reviewer reading at a normal pace will need to listen as well as read, because errors in names and numbers look plausible on the page.
Put the review where mistakes cost most: names, figures, dates and quoted statements. Sample a few minutes first and count errors; if the rate on your audio is higher than the headline number, plan more time before you start.
- Names, numbers and dates first.
- Sample a few minutes to get your own rate.
- Listen, do not only read.
Sources
Related posts
More in Use cases
- Meta Feed video ad spec checklist before uploading an AI clip
Meta's Facebook Feed video spec: 4:5, 1440 x 1800, up to 4 GB, MP4 or MOV, H.264 with AAC audio. A checklist, and the Sume Timeline output size that meets it.
- Meta Feed allows 1-second video ads: shortest Sume clip is 2 s
Meta's Feed spec allows video from 1 second. The shortest clip any Sume video model generates is 2 seconds (Wan 3.0). What to do for a 1-second cut test.
- Microscope-style science clip with Gemini Omni Flash: prompt and cost
Google's own Omni 1.1 demo prompt is a diatom micrograph. Here is how to write one, draft it at 360p, and what 8 seconds costs on Sume at each resolution.
- Midday to golden hour in a video: Omni edit or first/last frame
Change the time of day in a clip you already have. When a Gemini Omni Flash edit fits, when first and last frames fit, and what a 6-second pass costs on Sume.
Written by Sume