Index-Translate 2B, 9B or 35B-A3B for subtitle translation

Index-Translate ships 2B, 9B and 35B-A3B text models. What the vendor pages say about each size, and how to pick one for subtitle cues you burn with Sume.

5 min readSume
All posts

For subtitles, start with the hosted 35B-A3B endpoint if you translate a few clips, and move to the 9B or 2B weights only when you need to run translation on your own hardware. Index-Translate comes in three text sizes, 2B, 9B and a 35B-A3B preview, all under Apache-2.0, and Bilibili's own page serves the 35B-A3B for free, so the size question is really a hosting question.

Subtitle lines are short, which is why even the smaller sizes are plausible. The vendor publishes benchmark numbers only for the largest model in the pages we read, so treat any claim about the small ones as something to test on your own cues.

What do the vendor pages say about each size?

The GitHub README lists 2B, 9B and 35B-A3B (preview) text models for 150 languages. The 35B-A3B is a mixture-of-experts model: 35B parameters in total, about 3B active per token. Its FP8 checkpoint was published on October 3, 2026, validated on NVIDIA A100, and the card sets a maximum model length of 4096 tokens. The README gives default serving context as 32768 tokens for the 2B and 9B.

Index-Translate text sizes as published, read 2026-10-04
SizeHow you can run itContext noted by the vendor
2BOpen weights, Apache-2.032768 default serving
9BOpen weights, Apache-2.032768 default serving
35B-A3B previewOpen weights, FP8 and FP4 builds, plus a free hosted endpoint4096 max model length on the FP8 card
Vendor-reported scores, 35B-A3BFLORES 0.8794, WMT26 76.76, instruction-following 0.8336Vendor figures, not independently checked

Which size fits which subtitle job?

The 4096-token limit on the FP8 card is plenty for a 60-second clip, which has a few hundred words, but it matters if you paste a whole transcript into one request. Translate in chunks that match your caption windows instead.

  • A one-off video or a pilot: call the hosted 35B-A3B endpoint. No GPU, no setup, and the page states no price.
  • A recurring workflow with client footage: self-host a size you can serve, so transcripts never leave your infrastructure.
  • A laptop or a small GPU: try the 2B first and compare it with the 9B on 20 real cues.
  • Many languages in one batch: keep one size across languages so the style stays consistent.

Where does Sume come in?

The model only produces text. Sume's video captions endpoint takes the translated lines as cues with start and end, burns them onto a public HTTPS video and skips speech-to-text. A request carries at most 200 cues of up to 400 characters each, inside a 60-second window, and a standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate. Switching from 35B-A3B to 9B changes nothing on the Sume side.

Because the output is identical in shape, you can run the same 20 cues through two sizes, burn both onto the same clip and compare them on screen before committing. The full request is in the cue-translation walkthrough.

How do you decide with evidence?

Do not choose on parameter count. Take 20 cues from your real video, including the longest line and a line with a brand name, and run each size you can serve over them. Compare line breaks, dropped words and how often a term is mangled, then burn the top two onto the same clip. Size matters less for short caption lines than for long documents, but only your own cues will show whether that holds for your language pair.

Record which size and which prompt produced which cues, so a later quality complaint can be traced to a setting rather than guessed at.

Sources

Related posts

More in Models

All Models posts

Written by Sume