AI hug video generator: from two photos to one hug clip
Make an AI hug video in two steps: combine both people in one still and check it, then animate it with a prompt that describes the hug.

An AI hug video generator turns photos of two people into a short clip of them hugging. It takes two steps: first combine both people into one still with an image model and check it, then animate that still as the first frame of an image-to-video clip, with the hug described in the prompt. Sending both photos straight to a video model as reference images is the one-step alternative, but references only guide the picture, so faces can drift.
Facts come from Sume's Image API, Video generation, and Media inputs docs, read on 2026-09-28. Anything described as current behavior is read from Sume's API code. Without code, describe the clip in the Agents tab: the agent picks the models and asks before it spends.
How do I make an AI hug video from two photos?
- Combine the two people in one still with
POST /v1/images: both photos ininput_references, and a prompt that places them facing each other in one setting. How to combine two photos into one with AI covers that call and which image models take two photos. - Check the still: both faces, both sets of hands, the setting. Generate again until it is right, because it becomes the clip's first frame.
- Host the still you keep at a public HTTPS URL. The Image API's result URLs are signed, and signed or private URLs are refused as generation inputs.
- Animate it: send the still to
POST /v1/videosinframe_imagesas thefirst_frame, with the hug inprompt.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hug-clip-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "The two people turn toward each other, step closer and share a warm hug, holding it for a moment before they lean back and smile. Static medium shot, soft daylight",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/both-together.jpg" }, "frame_type": "first_frame" }
],
"aspect_ratio": "9:16",
"resolution": "720p",
"duration": 6
}'What should an AI hug video prompt say?
Leave the people's looks to the first frame and spend the prompt on what happens. Sume's docs suggest details about motion, camera angles, lighting, and scene composition:
- Who moves and how: "the woman on the left steps forward and wraps her arms around the man".
- How long it lasts: "they hold the hug for a moment, then lean back and smile".
- The camera: "static medium shot", so both people stay in frame.
- The light and mood: "soft window light" or "golden hour".
Can I skip the combined still?
Yes, with reference-to-video: send both photos in input_references as type: "image_url" entries and describe the hug in the prompt. References are visual guidance rather than exact frames, so nothing pins either face. Leave frame_images out of that request: frames take precedence, and in current code the references are then not sent to the model.
On Gemini Omni Flash 1.1, the prompt can point at each photo by position, <IMAGE_REF_0> for the first and <IMAGE_REF_1> for the second, as Gemini Omni Flash 1.1 video API explains. Reference-to-video API lists how many images each model takes.
| Still first | Photos as references | |
|---|---|---|
| Calls | POST /v1/images, then POST /v1/videos | POST /v1/videos |
| Video field | frame_images with frame_type: "first_frame" | input_references with type: "image_url" |
| What is pinned | The clip's first frame, which you approved | Nothing: visual guidance rather than exact frames |
| Models | Any model whose supported_frame_images lists first_frame | All except kling-3 and grok-imagine-video-1.5, which take no references |
Will the people look like themselves?
Not guaranteed. The first frame is your approved still, but every later frame is generated, and faces can change as people turn and move into the hug. Watch the clip and generate again if a face drifts. Use photos of people who agreed to it, or photos you have permission to use.
If one photo is an old print, scan it first; Can AI animate old photos? covers scans and what to expect.
What are the limits?
- One clip runs 2–30 seconds depending on the model.
- Photo, still, and frame URLs must be public HTTPS.
- Two steps, two bills: image generation is all-or-nothing, billed only when an image completes, and clips are billed by model at the rates
GET /v1/videos/modelslists inpricing_skus.
Sources
Related posts
More in Use cases
- AI icon generator: app icon art at 1024 × 1024 from a prompt
Generate app icon art from a prompt: one centered symbol, no text, several options per call, at 1024 × 1024. Then resize, cut out and finish it.
- AI image generator for presentations: 16:9 slide images
Make AI images for a presentation at the slide's shape, 16:9 by default in PowerPoint: exact pixel sizes, one style on every slide, room for text.
- What is AI info on Facebook and Instagram? Meta's AI label
AI info is Meta's label for AI-generated or AI-edited images, video and audio, added when Meta detects AI indicators or when the poster discloses AI.
- AI infographic generator: draw the layout, add the numbers
An image model can draw an infographic's layout, icons, and colors, but not reliable numbers. Generate a tall frame, then add the data yourself.
Written by Sume