Swap in an AI-generated presenter: portrait first, then Recast

No photo of a real person? Generate a presenter portrait with Sume's image API, pass its URL to h3-max-recast, and poll the video job. Python included.

5 min readSume
All posts

H3 Max Recast needs one photo per person it puts into the clip. fal's schema says the references replace "the main people in the video from left to right" (read 2026-10-04). Nothing in that rule says the photo has to show a real individual. If you want a presenter who is not a real person, and so no release to chase, generate the portrait first and hand its URL to Recast.

Both calls run on Sume. The portrait is an Image API request; the swap is a Videos API job.

The script

POST /v1/images waits up to 30 seconds and answers 200 with data[].url; a 202 means you got a job envelope instead, so the script stops there. The video submit sends the source clip and the portrait as input_references, with duration set to the source length in seconds. Replace the source URL with a public HTTPS clip of 5 to 30 seconds.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

img = requests.post(f"{API}/v1/images", headers=H, timeout=60, json={
    "model": "openai/gpt-image-2.5",
    "prompt": "Head-and-shoulders portrait of an adult presenter facing the camera, soft studio light, plain background",
})
if img.status_code != 200:
    raise SystemExit(img.text)
portrait = img.json()["data"][0]["url"]

job = requests.post(f"{API}/v1/videos", timeout=60,
    headers={**H, "Idempotency-Key": "recast-ai-presenter-001"}, json={
    "model": "h3-max-recast",
    "duration": 12,
    "input_references": [
        {"type": "video_url", "video_url": {"url": "https://example.com/source.mp4"}},
        {"type": "image_url", "image_url": {"url": portrait}},
    ],
})
job.raise_for_status()
poll = job.json()["polling_url"]
while True:
    body = requests.get(poll, headers=H, timeout=60).json()
    if body["status"] in ("completed", "failed", "cancelled"):
        break
    time.sleep(10)
print(body["status"], body.get("unsigned_urls"), body.get("usage"))

What it costs

The video side is $0.375 a second at the default 768p on Sume and $0.5625 at 1080p, so the 12 second job above is $4.50. The portrait is one image call whose billed cost comes back in usage.cost. Read the video's usage.cost too, and log both under one idempotency prefix.

What each call needs (Sume and fal docs, read 2026-10-04)
CallRequiredResult
POST /v1/imagesmodel, promptdata[].url, usage.cost
POST /v1/videos with h3-max-recastduration, one video and 1 to 4 images in input_referencespolling_url, then unsigned_urls

Two cautions

A generated face can still resemble someone real, so look at the portrait before it goes into an ad, and keep the image file with the job id. Second, one portrait gives one presenter: for a recurring spokesperson, keep the same portrait URL for every clip rather than regenerating it, because each call produces a new face.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume