Photo to talking selfie clip: put the spoken line in the Omni prompt
fal's Omni 1.1 example ends its prompt with a quoted spoken line. Here is that pattern on Sume from a still, the cost of 6 seconds, and a transcript check.

To get a short talking selfie from a photo with Gemini Omni Flash 1.1, send the photo as image_url and write the spoken words inside the prompt, after the scene description, in the form He is saying: "...". A 6-second 9:16 clip at 720p costs $0.75 on Sume. Then run the clip through video inspect with transcribe: true to see what was actually said.
Where the pattern comes from
The fal model page's example prompt describes a 9:16 smartphone selfie in a cafe in about a paragraph: lens and arm's-length framing, the face and clothing, the light from a window, the background, a statement about real skin texture, and then, as the last sentence, He is saying: "Third one today. Don't tell my doctor." (read 2026-10-07). The page lists the same endpoint's rates and says a 10-second 1080p clip costs $1.50 at fal list. Sume's rate is that list times 1.25.
The speech is generated with the picture. Audio is always on for this model, and there is no audio reference input, so you cannot hand it a recorded voice.
The request
Image-to-video takes image_url and an optional end_image_url. Duration is 3 to 10 seconds; for one short sentence, 5 or 6 seconds leaves room to speak without padding.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: selfie-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"image_url": "https://example.com/portrait.png",
"prompt": "Vertical smartphone selfie, arm held out, soft window light, natural skin texture. He is saying: \"Third one today.\"",
"duration": 6,
"aspect_ratio": "9:16",
"resolution": "720p",
"mode": "async"
}'Keep the line short
One short sentence fits a 6-second clip. A paragraph does not, and a line that is too long for the clip will not fit. Write the words the way people say them, with the punctuation that would make a pause, and avoid numbers and abbreviations that can be read several ways.
| Length | 720p ($0.125/s) | 1080p ($0.1875/s) |
|---|---|---|
| 5 seconds | $0.625 | $0.9375 |
| 6 seconds | $0.75 | $1.125 |
| 8 seconds | $1.00 | $1.50 |
Framing the still
The image sets the face, the framing and the light, so choose a photo that already looks like the clip you want: a face that fills the upper half of the frame, a mouth that is visible and not covered by a hand, eyes toward the lens. A photo of someone in profile will not become a convincing front-facing selfie, and a very small face in a wide shot gives the model little to animate. Use a high-resolution image, since Google's docs recommend high-resolution images for image-to-video.
Write the setting into the prompt too, even if it is visible in the photo. The prompt in fal's example restates the clothing, the glasses and the window light, and it also says what to avoid, such as smoothing or de-aging the face.
Verify the words
Import the finished clip, then call POST /v1/video-inspect with transcribe: true. The result carries a transcript with text and word timings. Compare the text with your script word for word before the clip goes out. Spoken-line prompts are a request, not a guarantee, so check names and numbers with extra care. For a client script that must be exact, the lip-sync route with fresh TTS audio is the safer path, and the Kling and Seedance transcript check shows the same verification on other models.
Sources
Related posts
More in Use cases
- Pick clip lengths from a voice-over script: one sentence per clip
Turn a voice-over into clips: measure each voiced sentence, round up, and pick the models whose duration window contains it. Windows for six Sume models.
- How do I turn one podcast episode into five quote clips for social?
Transcribe the episode in 10-minute chunks, split five quotes out, put the cover still under each and burn captions: $1.80 for a 28-minute episode on Sume.
- Post-call recap video with an AI avatar for prospects: script and cost
After a sales call, send a 30-second recap clip from a Sume avatar. Script structure, cost at standard, plus and max, and how to keep it honest and reviewed.
- Pre-rendered avatar greetings per visitor segment, not a live avatar
Instead of a live avatar for each visitor, render one short Sume avatar clip per segment ahead of time. Per-tier cost for six 12-second greetings.
Written by Sume