Add captions to a MiniMax H3 clip: video-captions, silent clips
A minimax-h3 clip carries native stereo audio, so video-captions can transcribe it. Silent clips fail with caption_no_speech; pass cues instead.

To caption a MiniMax H3 clip on Sume, pass the finished clip's URL to POST /v1/video-captions. H3 generates native stereo audio with no toggle, so speech-to-captions has something to transcribe if the clip contains speech. A clip with no audible speech fails as caption_no_speech; send authored cues with text, start and end to burn overlay copy without speech recognition.
Captions behavior comes from Sume's Video captions docs; H3's 5-15 s range and native audio come from the Video Router docs, read 2026-10-02. OrcaRouter reports H3 was announced 2026-07-31 with stereo audio (search summary only, read 2026-10-02).
Will the captions come from the H3 audio?
If you send no script_text, Sume runs speech-to-text on the clip. That needs audible speech; music or ambience alone does not count. If your prompt asked for dialogue, check the clip first, because a generated line may not match the words you wrote.
When you know the exact words, pass script_text for alignment, or pass cues to skip speech recognition altogether.
| Input | Effect |
|---|---|
video_url | Required; a public HTTPS clip URL |
script_text | Align your own text to the speech |
cues or segments | Authored text, start, end; skips speech-to-text |
language | Speech-to-text hint such as en or ko |
style | slam, punch, tiktok-green, korean-ad and others |
What happens with a silent H3 clip?
The job fails as caption_no_speech, with next_action: use_overlay_captions. It is not a policy rejection. Retry with cues carrying your lines and timings; the captions are burned onto the existing video, so the clip is not regenerated.
Keep the clean clip. Captions are a separate pass, and a caption failure leaves the original untouched, so you can recaption with a different style without paying for another H3 run.
When should I caption after generating, not during?
Generated on-screen text is unreliable, so keep text out of the prompt and add it afterwards.
H3 clips run 5 to 15 seconds, so one caption job covers a whole clip. For a longer piece, join clips first with Timeline 1.0, then caption once; the post on captions for long video covers limits.
A simple routine: generate the H3 clip, run video-inspect or listen to it, then caption. If the speech is correct, pass no text and let speech-to-text do the work; if it is not, pass cues with your own lines and timings and skip recognition.
Sources
Related posts
More in Media tools
- Pull the audio out of a MiniMax H3 clip with Sume audio-detach
A minimax-h3 clip has native stereo audio. audio-detach returns it as a wav or mp3 artifact without touching the video. Fields, defaults and caps.
- MiniMax H3 reference formats: HEIC, MOV, WAV, MP3 limits
MiniMax H3 accepts HEIC/HEIF images, MP4/MOV video and WAV/MP3 audio as references. A pre-flight checklist before you submit minimax-h3 on Sume.
- Keep the sound of AI clips in a Sume timeline: detach to a spine
A Timeline 1.0 render takes one audio spine. To keep audio generated with your clips, detach each clip's track, join the parts, and use that as the spine.
- Pika SFX: 1-20 second clips, negative prompts and seed vs Sume
Pika SFX makes 1 to 20 second effects with negative prompts, seed and guidance. Sume lists no sound-effects model, so here is what to use for each need.
Written by Sume