Graduation announcement video: first frame, then names as caption cues
Start a graduation clip from a cap-and-gown photo as first_frame on /v1/videos, then put the names and date on as caption cues instead of generated text.

For a graduation announcement, send the cap-and-gown photo as a first_frame entry in frame_images on POST /v1/videos, and add the graduate's name, school and date as caption cues with POST /v1/video-captions afterward. Cues are authored overlay text, so nothing depends on a video model spelling a name correctly.
Sources: Video generation and Video captions, read 2026-10-04.
Check the model first
frame_images entries need a frame_type of first_frame or last_frame. The docs' own example uses seedance-2; list GET /v1/videos/models and read supported_frame_images before using another model.
| Field | Value | Note |
|---|---|---|
| model | seedance-2 | The documented first_frame example; read supported_frame_images for others |
| frame_images | One entry, frame_type first_frame | image_url.url is a public HTTPS image |
| resolution | 720p or 1080p | Docs examples use both |
| aspect_ratio | 9:16 or 16:9 | Match your photo |
Generate
Ask for small motion that starts from the photo: a cap toss, confetti, a slow push-in. Because the photo is the first frame, the graduate's face should match it at the start.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "The graduate laughs and tosses the cap in the air, confetti falls, slow push-in. Keep the gown color unchanged.",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/graduate.jpg" }, "frame_type": "first_frame" }
],
"resolution": "1080p",
"aspect_ratio": "9:16"
}\Fetch the clip
Poll polling_url until the status is completed, then fetch the video from unsigned_urls[0]. Import it to media.sume.com before it goes to any other Sume job, since the media routes read only workspace URLs.
Add the names as cues
The cue form is text, start, end in seconds. It skips speech-to-text, which is what you want for a clip that may be silent, because a silent clip with no cues fails as caption_no_speech.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: grad-cues-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/graduate.mp4",
"style": "slam",
"cues": [
{ "text": "Maya Chen", "start": 0.5, "end": 3.0 },
{ "text": "Class of 2026", "start": 3.0, "end": 5.5 },
{ "text": "Lincoln High School", "start": 5.5, "end": 8.0 }
]
}\Sources
Related posts
More in Use cases
- Green Monday 2026 ad plan: six product clips and captions for $3.18
Green Monday falls on Dec 14, 2026, 71 days after today. A six-clip Wan 3.0 plan with captions and one Timeline render costs $3.18 on Sume; the maths inside.
- Halloween video sound design: a Lyria prompt that stays scary
Halloween trailer and ad sound design with one prompt to Sume's Music Router: drones, stingers and a quiet bed, no vocals, plus a silent-clip caption fallback.
- Haunted house ticket teaser: three stills, one 18 s render, $2.68
A haunted house or trail promo from three photos: Wan 3.0 clips at $0.125 a second, a Timeline render, a music bed and burned-in dates. $2.675 in total.
- HeyGen Free: 3 videos of 1 minute to test a script
HeyGen's Free plan lists 3 videos a month up to a minute. Use them to judge a script's pacing, then check it against Sume's 4 to 60 second window.
Written by Sume