Index-Translate, Echo, Homura: which part do you need?

Bilibili's Index-Translate release has text, speech, dubbing and length parts. Which fit subtitles, and which Sume step follows.

5 min readSume
All posts

For subtitles, you need one of two parts: the text model if you already have a transcript, or Index-Echo S2TT if you only have a video with Chinese speech. The Index-Translate release of September 30, 2026 is a family, and the README names several pieces that sound alike. Here is which one does what, and which are irrelevant to burned captions.

All the pieces below were read on the vendor's pages on 2026-10-04, and we have run none of them.

What is each part for?

Pick by input and output. The Sume step in the last column is what follows when you want captions burned onto a video.

Index-Translate family parts, read 2026-10-04
PartJob as described by the vendorNext Sume step
Index-Translate (2B, 9B, 35B-A3B preview)Text translation across 150 languagesSend translated cues to video captions
Index-Echo S2TTSpeech to bilingual SRT with timestamps, in 60-second windowsConvert the SRT to cues, then burn
Index-Echo S2STSpeech-to-speech dubbingNo documented Sume step; audio goes to your editor
Index-HomuraSyllable-controlled translationBudget script length before voice
Index-NativeLongLong-text translationCut into 60-second cue windows for captions

Which part for which task?

We have not checked how the long-text part behaves with subtitles, so for captions treat any long source as a set of short windows regardless.

  • You have a transcript or SRT: the text model, then cue validation.
  • You have only a Chinese-speech video: Index-Echo S2TT, as in the S2TT post.
  • You need the voice dubbed: S2ST, then mix and join the audio yourself.
  • You need translated speech to fit a time slot: Homura, plus a length check, as in the length post.

What does Sume add?

None of these models burn captions onto video. Sume's video captions endpoint takes your cues, up to 200 per request and 400 characters each within 60 seconds, and renders them on a public HTTPS video. A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate.

Which size of the text model to start with is covered in the size guide. The sizes and the hosted endpoint are separate from this map of parts.

What should you not assume?

Three cautions apply before you build on any of these. First, the 35B-A3B text model is labelled a preview, so its behaviour may change. Second, the hosted endpoint is described as an online demo, which is not a service commitment. Third, the speech parts are packaged for a small set of directions, with Chinese to English, Japanese and Spanish named on the S2TT card, so other pairs need the text model plus your own speech-to-text.

If a limit matters to your project, read the current model card again on the day you build, since these pages were fresh on 2026-10-04 and are still moving.

Sources

Related posts

More in Models

All Models posts

Written by Sume