Add captions to a Vidu Q4 clip with Sume: what the caption API needs
Vidu Q4 Preview can voice characters. Sume's video captions API burns captions onto any public HTTPS clip; it fails on silent clips unless you send cues.

The short answer
Yes, you can caption a clip from Vidu Q4 Preview on Sume, even though Sume does not list Vidu as a model. Sume's video captions API takes any public HTTPS video URL, so a finished Vidu clip hosted at a fetchable address works as input. Send POST /v1/video-captions with video_url. If the clip has speech, captions come from speech-to-text. If it is silent, the job fails with caption_no_speech unless you send your own cues.
What the request needs
The launch release says Vidu Q4 Preview adds reference voice, so many clips will carry dialogue. The table lists the fields from the Sume docs that matter for a generated clip.
| Field | Required | Effect |
|---|---|---|
| video_url | Yes | Public HTTPS URL of the finished clip |
| style | No | slam, punch, tiktok-green, korean-ad and others; the caption text sets a default if you omit it |
| language | No | Hint for speech-to-text only, such as en or ko |
| script_text | No | Your known dialogue, aligned to the speech |
| cues or segments | No | Authored text with start and end, used with no speech-to-text |
| webhook_url | No | Receive a notification when the job finishes |
Silent and scripted clips
Generated video is often silent, or its audio is music. The caption job needs audible speech; otherwise it returns caption_no_speech with a hint to use overlay captions. For those clips send cues, each with text, start and end in seconds, and Sume burns the text with no transcription step.
If you wrote the dialogue yourself, pass it as script_text. The docs describe this as aligning your script to the speech, which fixes misheard names, product terms and brand spellings that speech-to-text gets wrong.
Getting the URL right
The input must be a fetchable public HTTPS URL; the API rejects localhost, private-network and non-HTTPS addresses before the job starts. If a vendor serves results from a link that expires, copy the file to storage you control first, then caption from there. The asset library docs list the input rules for each route. Look at the captions docs for the style list and the design overrides. Test one style on a short clip first, since the default for Latin text is a large single-word style that suits vertical video more than a wide cinematic frame.
Order of operations
Burned captions are pixels in the frame, so caption last. Trim, cut and assemble first, then caption the finished clip, so the timings match the final edit. Keep an uncaptioned copy too, because you cannot remove burned text later, and a platform that renders its own captions will show both. The docs example uses a file called clean.mp4 for exactly this reason. If you need the same clip in two languages, run two caption jobs from the one clean source and change language or script_text between them.
Sources
Related posts
More in Media tools
- Caption 30 product videos with Sume Video captions for $6
Sume Video captions reserve and capture $0.20 per standalone job, so 30 product videos cost $6.00. Per-batch arithmetic and what the job returns.
- Caption a 45-second reel for 20 cents: script_text or speech-to-text
A standalone Sume caption job is $0.20 for a clip up to 60 seconds, so a 45-second reel is 20 cents. Send script_text to keep your exact wording.
- Caption a 75-second clip: the $0.20 estimate is for 60 seconds
Sume's standalone caption job lists $0.20 for videos up to 60 seconds. For a 75-second clip, read the live catalog price or split the cut first. Budget table.
- caption_no_speech: why a caption job fails on a silent clip
A Sume caption job returns caption_no_speech when the clip has no audible speech. The fix is cues, segments or words, not a retry.
Written by Sume