Graduation announcement video: first frame, then names as caption cues

Start a graduation clip from a cap-and-gown photo as first_frame on /v1/videos, then put the names and date on as caption cues instead of generated text.

5 min readSume
All posts

For a graduation announcement, send the cap-and-gown photo as a first_frame entry in frame_images on POST /v1/videos, and add the graduate's name, school and date as caption cues with POST /v1/video-captions afterward. Cues are authored overlay text, so nothing depends on a video model spelling a name correctly.

Sources: Video generation and Video captions, read 2026-10-04.

Check the model first

frame_images entries need a frame_type of first_frame or last_frame. The docs' own example uses seedance-2; list GET /v1/videos/models and read supported_frame_images before using another model.

Fields for the first-frame clip (read 2026-10-04)
FieldValueNote
modelseedance-2The documented first_frame example; read supported_frame_images for others
frame_imagesOne entry, frame_type first_frameimage_url.url is a public HTTPS image
resolution720p or 1080pDocs examples use both
aspect_ratio9:16 or 16:9Match your photo

Generate

Ask for small motion that starts from the photo: a cap toss, confetti, a slow push-in. Because the photo is the first frame, the graduate's face should match it at the start.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "seedance-2",
    "prompt": "The graduate laughs and tosses the cap in the air, confetti falls, slow push-in. Keep the gown color unchanged.",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/graduate.jpg" }, "frame_type": "first_frame" }
    ],
    "resolution": "1080p",
    "aspect_ratio": "9:16"
  }\

Fetch the clip

Poll polling_url until the status is completed, then fetch the video from unsigned_urls[0]. Import it to media.sume.com before it goes to any other Sume job, since the media routes read only workspace URLs.

Add the names as cues

The cue form is text, start, end in seconds. It skips speech-to-text, which is what you want for a clip that may be silent, because a silent clip with no cues fails as caption_no_speech.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: grad-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/graduate.mp4",
    "style": "slam",
    "cues": [
      { "text": "Maya Chen", "start": 0.5, "end": 3.0 },
      { "text": "Class of 2026", "start": 3.0, "end": 5.5 },
      { "text": "Lincoln High School", "start": 5.5, "end": 8.0 }
    ]
  }\

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume