Pick a subtitle translation model with a 20-cue pilot

Vendor benchmark scores do not tell you how a model handles your captions. Run the same 20 cues through each candidate, burn them, and compare on screen.

5 min readSume
All posts

Run the same 20 cues from your own video through each candidate model, score them with a few mechanical checks, and burn the best two onto the clip to compare on screen. Published scores are a starting list, not a verdict. The Index-Translate README reports FLORES 0.8794, WMT26 76.76 and an instruction-following score of 0.8336 for its 35B-A3B preview (vendor-reported, read 2026-10-04), and none of those measures caption length, line breaks or your brand terms.

Twenty cues is enough to expose the common failures and small enough to read by eye.

Which 20 cues?

Do not take the first twenty. Choose the ones most likely to break a model, because easy lines will make every candidate look the same.

A 20-cue pilot set, read 2026-10-04
Cue typeHow manyWhat it tests
Shortest lines, one or two words4Context-free fragments
Longest lines, near 400 characters3Length and line breaks
Brand and product names4Glossary compliance
Numbers, prices, dates3Number formats
Slang, idiom, jokes3Register and tone
Ordinary sentences3A baseline

What do you score?

The first three are code and cost nothing. The glossary post and the reading-speed post have helpers for them. Reserve human time for the fourth.

  • Count and timing preserved, with no merged or dropped cues.
  • Protected terms kept verbatim.
  • Length against the time slot, so the line can be read.
  • A native reader's pass on meaning, if you can get one.

How do you compare on screen?

Post each candidate's cues to Sume's video captions endpoint on the same short public HTTPS clip with the same style. A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate, so two finalists cost about forty cents, which is cheap next to burning the wrong choice into a whole library.

Cues are capped at 200 per request, so a 20-cue pilot is well inside the limit. Then choose by what you see: line breaks, reading speed and the brand names are all visible in the burned clip in a way a score is not. Model sizes to try first are in the size guide.

When is a pilot not enough?

A 20-cue pilot is a screen, not a proof. It will catch obvious failures such as dropped cues, mangled brand names and lines that cannot be read in time. It will not tell you how the model behaves across an hour of content, or in a rare dialect. For high-stakes work, add a human reviewer for the final pass, and keep a small sample of every batch to spot-check.

If two candidates tie, choose the one you can serve more cheaply or more privately, since the quality difference is then too small to matter.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume