Swap in an AI-generated presenter: portrait first, then Recast
No photo of a real person? Generate a presenter portrait with Sume's image API, pass its URL to h3-max-recast, and poll the video job. Python included.

H3 Max Recast needs one photo per person it puts into the clip. fal's schema says the references replace "the main people in the video from left to right" (read 2026-10-04). Nothing in that rule says the photo has to show a real individual. If you want a presenter who is not a real person, and so no release to chase, generate the portrait first and hand its URL to Recast.
Both calls run on Sume. The portrait is an Image API request; the swap is a Videos API job.
The script
POST /v1/images waits up to 30 seconds and answers 200 with data[].url; a 202 means you got a job envelope instead, so the script stops there. The video submit sends the source clip and the portrait as input_references, with duration set to the source length in seconds. Replace the source URL with a public HTTPS clip of 5 to 30 seconds.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
img = requests.post(f"{API}/v1/images", headers=H, timeout=60, json={
"model": "openai/gpt-image-2.5",
"prompt": "Head-and-shoulders portrait of an adult presenter facing the camera, soft studio light, plain background",
})
if img.status_code != 200:
raise SystemExit(img.text)
portrait = img.json()["data"][0]["url"]
job = requests.post(f"{API}/v1/videos", timeout=60,
headers={**H, "Idempotency-Key": "recast-ai-presenter-001"}, json={
"model": "h3-max-recast",
"duration": 12,
"input_references": [
{"type": "video_url", "video_url": {"url": "https://example.com/source.mp4"}},
{"type": "image_url", "image_url": {"url": portrait}},
],
})
job.raise_for_status()
poll = job.json()["polling_url"]
while True:
body = requests.get(poll, headers=H, timeout=60).json()
if body["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(10)
print(body["status"], body.get("unsigned_urls"), body.get("usage"))What it costs
The video side is $0.375 a second at the default 768p on Sume and $0.5625 at 1080p, so the 12 second job above is $4.50. The portrait is one image call whose billed cost comes back in usage.cost. Read the video's usage.cost too, and log both under one idempotency prefix.
| Call | Required | Result |
|---|---|---|
| POST /v1/images | model, prompt | data[].url, usage.cost |
| POST /v1/videos with h3-max-recast | duration, one video and 1 to 4 images in input_references | polling_url, then unsigned_urls |
Two cautions
A generated face can still resemble someone real, so look at the portrait before it goes into an ad, and keep the image file with the job id. Second, one portrait gives one presenter: for a recurring spokesperson, keep the same portrait URL for every clip rather than regenerating it, because each call produces a new face.
Sources
Related posts
More in Developers
- Find Sume jobs still queued or processing after 30 minutes
A scheduled script lists queued and processing jobs from GET /v1/jobs, picks those older than a cutoff, and prints each so you recover instead of resubmit.
- Symfony 7 controller for Sume webhooks: getContent and hash_equals
A Symfony controller that reads getContent(), recomputes hash_hmac over timestamp.raw_body, compares with hash_equals, and returns 204 for a Sume delivery.
- Synthesia needs a Legacy (v2) key; Sume keys carry scopes per call
Synthesia's quickstart warns an Interactive Avatars key will not create videos. Sume uses one bearer key whose scopes gate Formats, jobs and webhooks.
- Synthesia says 3 to 5 minutes per video; how to watch a Sume job
Synthesia's quickstart says videos usually finish in 3 to 5 minutes. Sume avatar jobs run async; read status, then the events endpoint for a slow render.
Written by Sume