AI hug video generator: from two photos to one hug clip

Make an AI hug video in two steps: combine both people in one still and check it, then animate it with a prompt that describes the hug.

5 min readSume
All posts

An AI hug video generator turns photos of two people into a short clip of them hugging. It takes two steps: first combine both people into one still with an image model and check it, then animate that still as the first frame of an image-to-video clip, with the hug described in the prompt. Sending both photos straight to a video model as reference images is the one-step alternative, but references only guide the picture, so faces can drift.

Facts come from Sume's Image API, Video generation, and Media inputs docs, read on 2026-09-28. Anything described as current behavior is read from Sume's API code. Without code, describe the clip in the Agents tab: the agent picks the models and asks before it spends.

How do I make an AI hug video from two photos?

  • Combine the two people in one still with POST /v1/images: both photos in input_references, and a prompt that places them facing each other in one setting. How to combine two photos into one with AI covers that call and which image models take two photos.
  • Check the still: both faces, both sets of hands, the setting. Generate again until it is right, because it becomes the clip's first frame.
  • Host the still you keep at a public HTTPS URL. The Image API's result URLs are signed, and signed or private URLs are refused as generation inputs.
  • Animate it: send the still to POST /v1/videos in frame_images as the first_frame, with the hug in prompt.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hug-clip-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "The two people turn toward each other, step closer and share a warm hug, holding it for a moment before they lean back and smile. Static medium shot, soft daylight",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/both-together.jpg" }, "frame_type": "first_frame" }
    ],
    "aspect_ratio": "9:16",
    "resolution": "720p",
    "duration": 6
  }'

What should an AI hug video prompt say?

Leave the people's looks to the first frame and spend the prompt on what happens. Sume's docs suggest details about motion, camera angles, lighting, and scene composition:

  • Who moves and how: "the woman on the left steps forward and wraps her arms around the man".
  • How long it lasts: "they hold the hug for a moment, then lean back and smile".
  • The camera: "static medium shot", so both people stay in frame.
  • The light and mood: "soft window light" or "golden hour".

Can I skip the combined still?

Yes, with reference-to-video: send both photos in input_references as type: "image_url" entries and describe the hug in the prompt. References are visual guidance rather than exact frames, so nothing pins either face. Leave frame_images out of that request: frames take precedence, and in current code the references are then not sent to the model.

On Gemini Omni Flash 1.1, the prompt can point at each photo by position, <IMAGE_REF_0> for the first and <IMAGE_REF_1> for the second, as Gemini Omni Flash 1.1 video API explains. Reference-to-video API lists how many images each model takes.

From Video generation, Video Router, Image API, and the catalog behind GET /v1/videos/models, read 2026-09-28.
Still firstPhotos as references
CallsPOST /v1/images, then POST /v1/videosPOST /v1/videos
Video fieldframe_images with frame_type: "first_frame"input_references with type: "image_url"
What is pinnedThe clip's first frame, which you approvedNothing: visual guidance rather than exact frames
ModelsAny model whose supported_frame_images lists first_frameAll except kling-3 and grok-imagine-video-1.5, which take no references

Will the people look like themselves?

Not guaranteed. The first frame is your approved still, but every later frame is generated, and faces can change as people turn and move into the hug. Watch the clip and generate again if a face drifts. Use photos of people who agreed to it, or photos you have permission to use.

If one photo is an old print, scan it first; Can AI animate old photos? covers scans and what to expect.

What are the limits?

  • One clip runs 2–30 seconds depending on the model.
  • Photo, still, and frame URLs must be public HTTPS.
  • Two steps, two bills: image generation is all-or-nothing, billed only when an image completes, and clips are billed by model at the rates GET /v1/videos/models lists in pricing_skus.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume