Quinceañera invitation teaser: portrait and venue as references
A 10-second quinceañera invitation teaser from up to 10 reference images with Gemini Omni Flash 1.1, with the date added as a plate instead of generated text.

For a quinceañera invitation teaser, send the honoree's portrait, the dress, the venue and the invitation art as reference_image_urls (up to 10) to gemini-omni-flash-1.1, and point at each one in the prompt as <IMAGE_REF_0>, <IMAGE_REF_1> and so on. You get a 3 to 10 second vertical clip with synced audio, and the date and address go on afterward as a plate or captions, not inside the generated video.
Facts are from the Video Router and Video generation pages, read 2026-10-04.
How the references work
Reference-to-video is one of four request shapes Gemini Omni Flash 1.1 routes by. The reference list is addressed by position, starting at 0.
| Input | Limit | How you address it |
|---|---|---|
| reference_image_urls | Up to 10 images | <IMAGE_REF_0>, <IMAGE_REF_1>, in list order |
| reference_video_urls | Up to 3 clips, each 3 s or shorter | <VIDEO_REF_0> and so on |
| reference_audio_urls | Not accepted | No audio references on this model |
| Duration | 3 to 10 seconds | 16:9 or 9:16 |
| Audio out | Always on | generate_audio false is rejected |
Submit the teaser
List the images in the order you will mention them. Four are enough for most invitations: a clear portrait first, then the dress, then the venue exterior, then the invitation card for color and type.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: quince-teaser-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A vertical invitation teaser. The young woman in <IMAGE_REF_0> turns toward camera in the dress from <IMAGE_REF_1>, in the garden courtyard from <IMAGE_REF_2>. Warm evening light, petals drifting, the color palette of <IMAGE_REF_3>.",
"reference_image_urls": [
"https://example.com/portrait.jpg",
"https://example.com/dress.jpg",
"https://example.com/venue.jpg",
"https://example.com/invitation-card.jpg"
],
"resolution": "1080p",
"aspect_ratio": "9:16",
"duration": 10,
"mode": "async"
}'Put the details on afterward
Do not ask the video model to render the date, time or address. Make the details as a still plate in your design tool, import it, and put it on the clip with Timeline compose in overlay mode (position, width_ratio 0.05 to 1, margin_ratio). Compose is a flat $0.02 job and the clip's length sets the output length.
If the invitation is bilingual, the cheapest safe route is the same: one plate per language, one compose job each, same teaser underneath.
Likeness note
Reference images are likeness inputs. Use photos of the honoree that her family has agreed to share and tell guests when a clip is AI-generated. Sume's docs do not describe a consent check, so that step is yours.
Sources
Related posts
More in Use cases
- Re-voice a recording: Sume STT then TTS as two jobs
Transcribe a recording with Sume STT, fix the text, then speak it with Sume TTS. Two jobs, two prices, and where the human edit goes.
- Real estate agent intro clip: headshot, logo and listing photos
Send a headshot, your logo and listing photos as reference_image_urls to Gemini Omni Flash 1.1 and address each as <IMAGE_REF_n> for a 3 to 10 second intro.
- Listing clip from photos with Seedance 2.5, and what to disclose
Nine listing photos can feed a Seedance 2.5 reference-to-video request on Sume. Price a 30 s clip, plan the camera path, label it AI-generated.
- Real estate price-reduction clip from one listing photo: $0.83 each
A 5-second vertical price-drop clip from the hero photo with the new price burned in as cues: $0.825 per listing, $9.90 for 12, on Sume from docs.
Written by Sume