AI wedding invitation video maker: art, motion, music, text
An AI wedding invitation video is your invitation art animated into short clips, set to music, with names, date and venue burned on as typed text.

Yes, an AI tool can make a wedding invitation video, but in parts: animate your invitation art or a photo of the couple into short clips, set them to a music track, then burn the names, date, venue and RSVP line on as text you typed. Don't ask the video model to write the names: every frame after the first is generated, so lettering it draws can drift, while typed captions come out spelled exactly as you typed them.
With Sume, you can describe the invitation in the Agents tab, where the agent picks the models and asks before it spends, or make the same four calls over the API. Facts below come from the Video generation, Music Router, Timeline 1.0 and Video captions docs, read on 2026-09-29; limits marked as current behavior are read from Sume's code.
How do I make a wedding invitation video with AI?
The pipeline is the same four calls as an AI birthday video from photos: animate stills into clips, generate a track, join them in one render, then burn the text on. What changes for a wedding:
- Start from a still you control: the invitation design, a save-the-date photo, or card art from an AI invitation generator with the text area left empty. It goes to
POST /v1/videosas thefirst_frameinframe_images, at a public HTTPS URL, with a prompt for small motion such as “petals drift down, candle flames flicker, slow push-in”. - Brief the music for the occasion: “Warm string quartet, 70 BPM, a gentle swell at 0:20. A 30-second track. Instrumental, no vocals.” Length is steered in the prompt; the Music Router rejects
duration. - Several events, one set of clips: a save-the-date, a mehndi or sangeet invitation and a reception invitation can reuse the same edit; each version is its own caption job with its own text.
How do I get the names and date on screen exactly right?
Burn them in as cues. Each cue is a text with a start and an end in seconds; cues skip speech-to-text and burn exactly that copy at those times. One card per detail keeps each line readable: names, then the date, then the venue, then the RSVP line.
Caption the finished edit, after the music is in. In current code the caption job refuses a video longer than 60 seconds or one with no audio stream, even when you send cues, so keep the invitation to a minute or less. Leave style out and Latin text gets slam, which in current code sets the words in capitals. How to add text over a video covers placement.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: invite-text-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/invite-edit.mp4",
"cues": [
{ "text": "Ana & Sam are getting married", "start": 0, "end": 5 },
{ "text": "Saturday, June 12", "start": 5, "end": 9 },
{ "text": "RSVP by May 1", "start": 20, "end": 25 }
]
}'Can the invitation be in Hindi or another script?
Not as burned captions, as far as the docs go. In the caption docs, font picks a face only for Hangul styles, the face list is Hangul, and any other font name is rejected rather than substituted. No Devanagari face is documented. For an Indian wedding invitation in Hindi, put the text into the invitation art yourself and use that still as the first frame; the opening frame keeps it as designed, but generated frames after it may not, so hold the text frames short or add motion only around them.
How much does an AI wedding invitation video cost?
Four calls, four prices. A 30-second invitation with three clips uses three video jobs, one track, one render and one caption job. Each extra event version over the same edit adds one caption job.
| Step | Call | Price |
|---|---|---|
| Animate the art | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Music | POST /v1/music-router/generate | $0.125 per audio |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Details as text | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What should I check before sending it?
- Faces: if you animate a photo of the couple, watch every clip. Only the first frame is your photo, so generate again if someone looks different.
- Frame shape: pick one
aspect_ratio, such as9:16for a phone message, for every clip. The render's default output is 1080×1920; setoutput.widthandoutput.heightfor landscape. - Length per clip: it depends on the model, and the longest accept 30 seconds; check
supported_durationsonGET /v1/videos/models. - Inputs: photo URLs must be public HTTPS. The render takes only this workspace's
media.sume.comfiles, such as the clips and track Sume returns, not files from your computer.
Sources
Related posts
More in Use cases
- AI whiteboard animation generator: blank board to sketch
AI can imitate whiteboard animation: a blank-board first frame, a finished-sketch last frame, the drawing in between. Words go on as captions.
- Background music for commercials: generate a bed that fits
Background music for a commercial is an instrumental bed under the voiceover that ends on time. How to brief, generate, cut and duck one with AI.
- Can I use AI images on Amazon KDP? What you must disclose
Yes, but KDP requires you to disclose AI-generated images, including cover and interior art. AI-assisted edits of your own work need no disclosure.
- Can I use AI images on Pinterest? What the AI label means
Yes. Pinterest labels Pins it detects as AI-generated or AI-edited "AI modified", from metadata or its classifiers, and you can appeal a wrong label.
Written by Sume