Add captions to a Vidu Q4 clip with Sume: what the caption API needs

Vidu Q4 Preview can voice characters. Sume's video captions API burns captions onto any public HTTPS clip; it fails on silent clips unless you send cues.

5 min readSume
All posts

The short answer

Yes, you can caption a clip from Vidu Q4 Preview on Sume, even though Sume does not list Vidu as a model. Sume's video captions API takes any public HTTPS video URL, so a finished Vidu clip hosted at a fetchable address works as input. Send POST /v1/video-captions with video_url. If the clip has speech, captions come from speech-to-text. If it is silent, the job fails with caption_no_speech unless you send your own cues.

What the request needs

The launch release says Vidu Q4 Preview adds reference voice, so many clips will carry dialogue. The table lists the fields from the Sume docs that matter for a generated clip.

Video captions request fields, as of 2026-10-08 (Sume docs, read 2026-10-08)
FieldRequiredEffect
video_urlYesPublic HTTPS URL of the finished clip
styleNoslam, punch, tiktok-green, korean-ad and others; the caption text sets a default if you omit it
languageNoHint for speech-to-text only, such as en or ko
script_textNoYour known dialogue, aligned to the speech
cues or segmentsNoAuthored text with start and end, used with no speech-to-text
webhook_urlNoReceive a notification when the job finishes

Silent and scripted clips

Generated video is often silent, or its audio is music. The caption job needs audible speech; otherwise it returns caption_no_speech with a hint to use overlay captions. For those clips send cues, each with text, start and end in seconds, and Sume burns the text with no transcription step.

If you wrote the dialogue yourself, pass it as script_text. The docs describe this as aligning your script to the speech, which fixes misheard names, product terms and brand spellings that speech-to-text gets wrong.

Getting the URL right

The input must be a fetchable public HTTPS URL; the API rejects localhost, private-network and non-HTTPS addresses before the job starts. If a vendor serves results from a link that expires, copy the file to storage you control first, then caption from there. The asset library docs list the input rules for each route. Look at the captions docs for the style list and the design overrides. Test one style on a short clip first, since the default for Latin text is a large single-word style that suits vertical video more than a wide cinematic frame.

Order of operations

Burned captions are pixels in the frame, so caption last. Trim, cut and assemble first, then caption the finished clip, so the timings match the final edit. Keep an uncaptioned copy too, because you cannot remove burned text later, and a platform that renders its own captions will show both. The docs example uses a file called clean.mp4 for exactly this reason. If you need the same clip in two languages, run two caption jobs from the one clean source and change language or script_text between them.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume