How do I caption a silent micro-drama for Shorts with overlay cues?
Send authored text as cues with start and end times to the Sume caption job. A silent clip fails speech captions with caption_no_speech, so use cues.

Use the cues field of POST /v1/video-captions with text, start and end for each line. Sume burns your text without speech-to-text, which is what a silent micro-drama needs.
If you send a silent clip without cues, speech captions fail with caption_no_speech and next_action: use_overlay_captions.
Why a micro-drama suits this
Short vertical dramas often run on music and expressions with title cards for dialogue. YouTube's Help page on Shorts series describes a series as a Shorts-only playlist with seasons and episodes, with each episode at most 180 seconds (read 2026-10-07). Fifteen to sixty seconds per episode is enough for a cliffhanger.
The caption job is $0.20 for videos of at most 60 seconds. A season of 30 silent episodes with authored cues is $6.00.
| Field | Meaning |
|---|---|
| cues (or segments) | Authored lines with text, start, end |
| style | slam, punch, tiktok-green, black-outline |
| design | Per-request color, type and motion overrides |
| Failure | caption_no_speech if no cues on a silent clip |
Writing the cues
Keep each line short enough to read in its window, and give the first line a start of zero so the hook is on screen at the first frame. Use one style for the whole series so episodes feel related.
Generated video with native audio is not silent, so set generate_audio to false on the video request if the model allows it and you want a clean music-only bed. The audio spine post shows how to add a theme. Details are on the captions page.
- First cue starts at 0.
- One idea per cue.
- Same
styleacross the season.
Pacing a silent drama
Silent episodes lean on the visual rhythm. Plan one new image every two or three seconds and one caption for each image, so the viewer always has something to read and something to look at. Fifteen seconds then holds five or six beats.
If you generate the beats with a video model that makes native audio, decide whether you want it. A silent drama usually wants music only. Set the audio option off in the request, or lay a music bed over the top in the Timeline render, as in the theme post.
Run one episode through the whole chain, from generated beats to burned cues, before you plan the season. At $0.20 for the cues you learn how the style looks on a phone, and how the first line reads, for a small spend.
A sample cue sheet
A 15-second episode might use four cues: 0.0 to 3.0 seconds for the hook line, 3.0 to 7.0 for the reveal, 7.0 to 11.0 for the reaction and 11.0 to 15.0 for the cliffhanger. Each text is under eight words.
Write the sheet before you generate, and generate the beats to match its timing.
Finally, test the cues at the smallest screen you expect. A caption that fits on a tablet can crowd the edge of a phone. The design overrides let you adjust placement and type size per request, and a wrong value fails with a 400 before you pay for a render.
Sources
Related posts
More in Use cases
- Skincare routine clip at 3:4: Seedance 2.5, $6.98 for 12 seconds
Seedance 2.5 has no 4:5 ratio on Sume. Render 3:4 at 834x1112 for $6.98 (12 s, 720p) and crop to 4:5 for the feed. Prices at 480p and 1080p too.
- Slow-motion water splash prompt for Gemini Omni: what 'slow' does
Prompt a slow-motion splash in Gemini Omni: how to word the speed, why a 5-second clip is enough, and what 360p, 720p and 1080p cost on Sume.
- Sneaker drop teaser in 6 seconds: Gemini Omni timecodes and cost
A 6-second vertical sneaker teaser built from a product photo and two timecoded beats in Gemini Omni on Sume, with 720p and 1080p prices and a Veo comparison.
- Solar roof before and after clip: Wan 3.0 first and last frame, 7 s
Send the bare roof as first_frame and the finished array as last_frame to wan-3.0. A 7-second clip costs $0.88 at 720p on Sume. Rules for matching the photos.
Written by Sume