Index-Translate, Echo, Homura: which part do you need?
Bilibili's Index-Translate release has text, speech, dubbing and length parts. Which fit subtitles, and which Sume step follows.

For subtitles, you need one of two parts: the text model if you already have a transcript, or Index-Echo S2TT if you only have a video with Chinese speech. The Index-Translate release of September 30, 2026 is a family, and the README names several pieces that sound alike. Here is which one does what, and which are irrelevant to burned captions.
All the pieces below were read on the vendor's pages on 2026-10-04, and we have run none of them.
What is each part for?
Pick by input and output. The Sume step in the last column is what follows when you want captions burned onto a video.
| Part | Job as described by the vendor | Next Sume step |
|---|---|---|
| Index-Translate (2B, 9B, 35B-A3B preview) | Text translation across 150 languages | Send translated cues to video captions |
| Index-Echo S2TT | Speech to bilingual SRT with timestamps, in 60-second windows | Convert the SRT to cues, then burn |
| Index-Echo S2ST | Speech-to-speech dubbing | No documented Sume step; audio goes to your editor |
| Index-Homura | Syllable-controlled translation | Budget script length before voice |
| Index-NativeLong | Long-text translation | Cut into 60-second cue windows for captions |
Which part for which task?
We have not checked how the long-text part behaves with subtitles, so for captions treat any long source as a set of short windows regardless.
- You have a transcript or SRT: the text model, then cue validation.
- You have only a Chinese-speech video: Index-Echo S2TT, as in the S2TT post.
- You need the voice dubbed: S2ST, then mix and join the audio yourself.
- You need translated speech to fit a time slot: Homura, plus a length check, as in the length post.
What does Sume add?
None of these models burn captions onto video. Sume's video captions endpoint takes your cues, up to 200 per request and 400 characters each within 60 seconds, and renders them on a public HTTPS video. A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate.
Which size of the text model to start with is covered in the size guide. The sizes and the hosted endpoint are separate from this map of parts.
What should you not assume?
Three cautions apply before you build on any of these. First, the 35B-A3B text model is labelled a preview, so its behaviour may change. Second, the hosted endpoint is described as an online demo, which is not a service commitment. Third, the speech parts are packaged for a small set of directions, with Chinese to English, Japanese and Spanish named on the S2TT card, so other pairs need the text model plus your own speech-to-text.
If a limit matters to your project, read the current model card again on the day you build, since these pages were fresh on 2026-10-04 and are still moving.
Sources
Related posts
More in Models
- Inworld TTS 2 style steering and the Sume voiceover path
What is reported about Inworld TTS 2 style steering and 100+ languages, and how to produce a voiceover with Sume's tts_create tool and join takes.
- Nano Banana 3: does it exist? Current ids
Google autocomplete suggests a Nano Banana 3, but suggestions are not releases. What autocomplete lists and which Nano Banana id Sume accepts.
- Kling 4.0 Flash 3-second minimum vs Sume's 4-second kling-3 floor
Kling 4.0 Flash starts at 3 seconds. Sume's kling-3 starts at 4. The models that can render a 3-second clip on Sume, and what a 3 s request costs.
- Kling 4.0 Flash is first-frame only; what Sume's Kling 3 accepts
Kling 4.0 Flash early access takes a first-frame image, 20 s, 720p, SDR. Sume ships kling-3: 4 to 15 s, 720p/1080p, priced per second. Limits side by side.
Written by Sume