Photo to talking selfie clip: put the spoken line in the Omni prompt

fal's Omni 1.1 example ends its prompt with a quoted spoken line. Here is that pattern on Sume from a still, the cost of 6 seconds, and a transcript check.

4 min readSume
All posts

To get a short talking selfie from a photo with Gemini Omni Flash 1.1, send the photo as image_url and write the spoken words inside the prompt, after the scene description, in the form He is saying: "...". A 6-second 9:16 clip at 720p costs $0.75 on Sume. Then run the clip through video inspect with transcribe: true to see what was actually said.

Where the pattern comes from

The fal model page's example prompt describes a 9:16 smartphone selfie in a cafe in about a paragraph: lens and arm's-length framing, the face and clothing, the light from a window, the background, a statement about real skin texture, and then, as the last sentence, He is saying: "Third one today. Don't tell my doctor." (read 2026-10-07). The page lists the same endpoint's rates and says a 10-second 1080p clip costs $1.50 at fal list. Sume's rate is that list times 1.25.

The speech is generated with the picture. Audio is always on for this model, and there is no audio reference input, so you cannot hand it a recorded voice.

The request

Image-to-video takes image_url and an optional end_image_url. Duration is 3 to 10 seconds; for one short sentence, 5 or 6 seconds leaves room to speak without padding.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: selfie-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "image_url": "https://example.com/portrait.png",
    "prompt": "Vertical smartphone selfie, arm held out, soft window light, natural skin texture. He is saying: \"Third one today.\"",
    "duration": 6,
    "aspect_ratio": "9:16",
    "resolution": "720p",
    "mode": "async"
  }'

Keep the line short

One short sentence fits a 6-second clip. A paragraph does not, and a line that is too long for the clip will not fit. Write the words the way people say them, with the punctuation that would make a pause, and avoid numbers and abbreviations that can be read several ways.

Cost of a talking selfie clip on Sume (read 2026-10-07)
Length720p ($0.125/s)1080p ($0.1875/s)
5 seconds$0.625$0.9375
6 seconds$0.75$1.125
8 seconds$1.00$1.50

Framing the still

The image sets the face, the framing and the light, so choose a photo that already looks like the clip you want: a face that fills the upper half of the frame, a mouth that is visible and not covered by a hand, eyes toward the lens. A photo of someone in profile will not become a convincing front-facing selfie, and a very small face in a wide shot gives the model little to animate. Use a high-resolution image, since Google's docs recommend high-resolution images for image-to-video.

Write the setting into the prompt too, even if it is visible in the photo. The prompt in fal's example restates the clothing, the glasses and the window light, and it also says what to avoid, such as smoothing or de-aging the face.

Verify the words

Import the finished clip, then call POST /v1/video-inspect with transcribe: true. The result carries a transcript with text and word timings. Compare the text with your script word for word before the clip goes out. Spoken-line prompts are a request, not a guarantee, so check names and numbers with extra care. For a client script that must be exact, the lip-sync route with fresh TTS audio is the safer path, and the Kling and Seedance transcript check shows the same verification on other models.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume