AI birthday video from photos: animate, add music and text

Make an AI birthday video by animating a few photos into short clips, joining them over an original instrumental, and burning your message on screen.

6 min readSume
All posts

To make an AI birthday video, animate a few photos of the birthday person into short clips with an image-to-video model, join the clips over an original instrumental track, and put your message on screen as text. Each photo becomes the first frame of its clip, so the video opens on pictures you chose, and the words you type are burned in as text instead of being generated.

With Sume, you can describe the video in the Agents tab and review it as it goes: the agent picks the models and asks before it spends. Over the API it is four calls, below. Facts come from the Video generation, Music Router, Timeline 1.0, and Video captions docs, read on 2026-09-28; limits marked as current behavior are read from Sume's code.

How do I make a birthday video from photos with AI?

  • Animate each photo: send it to POST /v1/videos as the first_frame in frame_images, at a public HTTPS URL, with a prompt for gentle motion.
  • Make the music: send a brief to POST /v1/music-router/generate, such as “Bright acoustic pop, 112 BPM, ukulele and handclaps, a big finish at 0:25. A 30-second track. Instrumental, no vocals.” Ask for an original instrumental in your own words rather than a named song.
  • Join the clips: one Timeline 1.0 render, shown below, with the clips as video[] slots, audio.mode: "silence" for the length, and the track as a looped soundtrack at gain_db: 0, as in an image slideshow but with moving clips. The clips and the track are Sume outputs, so their media.sume.com URLs qualify. In current code the clips' own sound is dropped, so the track is what you hear.
  • Add the message: send the render's video_url to POST /v1/video-captions with your words as timed cues (next sections).
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: birthday-edit-001" \
  -d '{
    "audio": { "mode": "silence", "duration_seconds": 15 },
    "soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/birthday-track.mp3", "gain_db": 0, "loop": true, "fade_out_seconds": 2 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/photo-1.mp4", "start": 0, "duration": 5 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/photo-2.mp4", "start": 5, "duration": 5, "transition": { "type": "fade", "duration": 0.5 } },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/photo-3.mp4", "start": 10, "duration": 5, "transition": { "type": "fade", "duration": 0.5 } }
    ]
  }'

What should an AI birthday video prompt say?

  • Small motion that fits the photo, with a simple camera: “She laughs and leans toward the cake as the candles flicker. Slow push-in, warm indoor light.”
  • No names, ages, or dates: put those in the text step, where you control the exact words.
  • One frame shape for every clip: aspect_ratio: "9:16" for a vertical video or "16:9" for a landscape one, with the photos cropped to match. The render's default output is 1080×1920, so set output.width and output.height for a landscape video.
  • Photos of people who agreed to it, or photos you have permission to use. Faces can change: every frame after the first is generated, so watch each clip and generate again if someone looks different.

How do I put the birthday message on screen?

Burn it in with a caption job. Each cue is a text with a start and an end in seconds; cues skip speech-to-text and burn that copy at those times. The same steps make a party invitation video: put the date, time, and place in the cues.

Caption the finished edit, after the music is in. In current code the caption job refuses a video longer than 60 seconds or one with no audio stream, even when you send cues. Leave style out and Latin text gets slam, which in current code shows your words in capitals and lays a light dark tint over the frame. How to add text over a video covers where the line sits.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: birthday-text-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/birthday-edit.mp4",
    "cues": [
      { "text": "Happy 7th birthday, Maya!", "start": 0, "end": 4 },
      { "text": "Love, Grandma and Grandpa", "start": 11, "end": 15 }
    ]
  }'

How much does an AI birthday video cost?

Four calls, four prices. A 15-second video like the example has three clips, one track, one render, and one caption job.

From Video generation, Music Router, Timeline 1.0, and Video captions, read 2026-09-28. Each rate is plus a 5.5% agent fee by default; see API pricing.
StepCallPrice
Animate each photoPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
MusicPOST /v1/music-router/generate$0.125 per audio generation
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
MessagePOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What are the limits?

  • The caption step takes videos of 60 seconds or less in current code, so keep the finished video under that or split it; add captions to a long video shows how.
  • One clip runs 2–30 seconds depending on the model. Make each clip a second longer than its slot, such as 6 seconds for the 5-second slots above: in current code a fade needs the clip it fades into to run past its slot by the fade's length, and otherwise renders as a hard cut.
  • Photo URLs must be public HTTPS; the render and caption steps use the media.sume.com files Sume returns.
  • The music's length is steered in the prompt, not set: the Music Router rejects duration. A looped soundtrack covers a short track.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume