Add captions to a MiniMax H3 clip: video-captions, silent clips

A minimax-h3 clip carries native stereo audio, so video-captions can transcribe it. Silent clips fail with caption_no_speech; pass cues instead.

4 min readSume
All posts

To caption a MiniMax H3 clip on Sume, pass the finished clip's URL to POST /v1/video-captions. H3 generates native stereo audio with no toggle, so speech-to-captions has something to transcribe if the clip contains speech. A clip with no audible speech fails as caption_no_speech; send authored cues with text, start and end to burn overlay copy without speech recognition.

Captions behavior comes from Sume's Video captions docs; H3's 5-15 s range and native audio come from the Video Router docs, read 2026-10-02. OrcaRouter reports H3 was announced 2026-07-31 with stereo audio (search summary only, read 2026-10-02).

Will the captions come from the H3 audio?

If you send no script_text, Sume runs speech-to-text on the clip. That needs audible speech; music or ambience alone does not count. If your prompt asked for dialogue, check the clip first, because a generated line may not match the words you wrote.

When you know the exact words, pass script_text for alignment, or pass cues to skip speech recognition altogether.

Caption inputs from Sume's Video captions docs, read 2026-10-02.
InputEffect
video_urlRequired; a public HTTPS clip URL
script_textAlign your own text to the speech
cues or segmentsAuthored text, start, end; skips speech-to-text
languageSpeech-to-text hint such as en or ko
styleslam, punch, tiktok-green, korean-ad and others

What happens with a silent H3 clip?

The job fails as caption_no_speech, with next_action: use_overlay_captions. It is not a policy rejection. Retry with cues carrying your lines and timings; the captions are burned onto the existing video, so the clip is not regenerated.

Keep the clean clip. Captions are a separate pass, and a caption failure leaves the original untouched, so you can recaption with a different style without paying for another H3 run.

When should I caption after generating, not during?

Generated on-screen text is unreliable, so keep text out of the prompt and add it afterwards.

H3 clips run 5 to 15 seconds, so one caption job covers a whole clip. For a longer piece, join clips first with Timeline 1.0, then caption once; the post on captions for long video covers limits.

A simple routine: generate the H3 clip, run video-inspect or listen to it, then caption. If the speech is correct, pass no text and let speech-to-text do the work; if it is not, pass cues with your own lines and timings and skip recognition.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume