Pick a subtitle translation model with a 20-cue pilot
Vendor benchmark scores do not tell you how a model handles your captions. Run the same 20 cues through each candidate, burn them, and compare on screen.

Run the same 20 cues from your own video through each candidate model, score them with a few mechanical checks, and burn the best two onto the clip to compare on screen. Published scores are a starting list, not a verdict. The Index-Translate README reports FLORES 0.8794, WMT26 76.76 and an instruction-following score of 0.8336 for its 35B-A3B preview (vendor-reported, read 2026-10-04), and none of those measures caption length, line breaks or your brand terms.
Twenty cues is enough to expose the common failures and small enough to read by eye.
Which 20 cues?
Do not take the first twenty. Choose the ones most likely to break a model, because easy lines will make every candidate look the same.
| Cue type | How many | What it tests |
|---|---|---|
| Shortest lines, one or two words | 4 | Context-free fragments |
| Longest lines, near 400 characters | 3 | Length and line breaks |
| Brand and product names | 4 | Glossary compliance |
| Numbers, prices, dates | 3 | Number formats |
| Slang, idiom, jokes | 3 | Register and tone |
| Ordinary sentences | 3 | A baseline |
What do you score?
The first three are code and cost nothing. The glossary post and the reading-speed post have helpers for them. Reserve human time for the fourth.
- Count and timing preserved, with no merged or dropped cues.
- Protected terms kept verbatim.
- Length against the time slot, so the line can be read.
- A native reader's pass on meaning, if you can get one.
How do you compare on screen?
Post each candidate's cues to Sume's video captions endpoint on the same short public HTTPS clip with the same style. A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate, so two finalists cost about forty cents, which is cheap next to burning the wrong choice into a whole library.
Cues are capped at 200 per request, so a 20-cue pilot is well inside the limit. Then choose by what you see: line breaks, reading speed and the brand names are all visible in the burned clip in a way a score is not. Model sizes to try first are in the size guide.
When is a pilot not enough?
A 20-cue pilot is a screen, not a proof. It will catch obvious failures such as dropped cues, mangled brand names and lines that cannot be read in time. It will not tell you how the model behaves across an hour of content, or in a rare dialect. For high-stakes work, add a human reviewer for the final pass, and keep a small sample of every batch to spot-check.
If two candidates tie, choose the one you can serve more cheaply or more privately, since the quality difference is then too small to matter.
Sources
Related posts
More in Developers
- Pin AI music model ids: music_v2_5, v6, lyria-3.5 and sume/music-auto
Which model id to pin for ElevenLabs Music, Suno and Google Lyria, and what Sume does with sume/music-auto. A config table and a rule for logging the engine.
- 1080x1350 sent as aspect_ratio: what Ideogram, Grok, Imagen get
Send pixels instead of a ratio and Sume snaps to the nearest native ratio. Tested table for 1080x1350, 1200x628 and 1500x500 across four model families.
- Poll Sume audio jobs at the recommended interval, not a fixed sleep
The Sume job status response tells you when to poll next and whether the job is terminal. Read those fields instead of hardcoding sleep(2).
- Ported Sora wrapper blocked until done? Sume sync stops at 30 s
A wrapper that blocks until the video is done will time out on Sume sync mode, which waits at most 30 seconds. Return the job id and poll, never resubmit.
Written by Sume