Swap the presenter in a product demo for each market with Recast
Plan one source demo and several Recast jobs, one per presenter photo: cost for five variants, idempotency keys, and the voice caveat, using documented limits.

You can adapt one product demo to several markets by recasting the presenter once per market: one source video, one job per presenter photo, each billed at the source length. At 15 seconds and 768p that is about $5.63 per variant on Sume, or $28.15 for five, before any 1080p finals (Video Router docs and fal, read 2026-10-03). The catch is the voice: Recast keeps the source audio, so every variant will speak with the same voice and words as the source.
That makes this plan right for demos where the audio is language-neutral, or music and sound effects with on-screen text, and wrong for ones where a local voice is the point.
Plan the fan-out
A fan-out is a loop over presenters, not a new workflow. Each job is independent, so a failed variant never blocks the other four. Give each job its own Idempotency-Key derived from the market and the source, so a retry replays the same job rather than starting a new one.
| Step | What happens | Cost at 768p |
|---|---|---|
| 1. Probe | Confirm 5 to 30 s and no shot over 15 s | None to Recast |
| 2. Draft | Recast one market at 768p and review tracking | $5.63 |
| 3. Fan out | Four more markets at 768p | $22.52 |
| 4. Finals | Re-run approved variants at 1080p ($8.44 each) | up to $42.20 for five |
| 5. Captions | Burn per-market captions on each output | Separate job per clip |
A fan-out sketch
The loop below submits one job per presenter photo and prints the job ids. It is a plain async script with a stable key per market. Replace the URLs with your own public HTTPS links.
import asyncio
import os
import httpx
SOURCE = "https://example.com/demo-15s.mp4"
PRESENTERS = {
"us": "https://example.com/presenters/us.jpg",
"de": "https://example.com/presenters/de.jpg",
"br": "https://example.com/presenters/br.jpg",
}
async def submit(c: httpx.AsyncClient, market: str, photo: str) -> None:
body = {
"model": "h3-max-recast",
"video_url": SOURCE,
"reference_image_urls": [photo],
"resolution": "768p",
"mode": "async",
}
r = await c.post("/v1/video-router/generate", json=body,
headers={"Idempotency-Key": f"demo-recast-{market}-v1"})
r.raise_for_status()
print(market, r.json())
async def main() -> None:
key = os.environ.get("SUME_API_KEY")
if not key:
raise SystemExit("set SUME_API_KEY")
async with httpx.AsyncClient(
base_url="https://api.sume.com",
headers={"Authorization": f"Bearer {key}"},
timeout=60,
) as c:
await asyncio.gather(*(submit(c, m, p) for m, p in PRESENTERS.items()))
asyncio.run(main())What to review on each variant
Review every variant, not a sample. The replaced presenter must track across every cut, the product must remain visible and unchanged, and the on-screen text must still read correctly. Sume's Recast docs promise that motion, camera, cuts and sound are kept; they do not promise that props and screen content are untouched, so check them frame by frame. A fast way to do that is to pull stills from each output at the same timestamps and compare them in a grid.
Also review the people. Use photos of presenters who have agreed to appear, and keep their consent records with the campaign. Sume documents the tool and its limits; it does not tell you whether a given use is allowed in a given market.
- Validate with one draft before fanning out.
- Keep keys stable per market and source version.
- Hold finals at 1080p until the draft is approved.
What to tell reviewers before they see the variants
Variants fail review for boring reasons. Tell the reviewers what is expected to differ: the presenter's face and appearance, nothing else. Give them the source next to each output, and ask them to flag anything that changed beyond the person, such as lighting on the hands, a product label or text on a screen. That keeps feedback specific enough to act on, and it tells you early if a particular source is a poor fit for Recast, for example one with heavy occlusion or a presenter who leaves and re-enters the frame.
Keep a short log per market: the source version, the photo used, the resolution, the job id and the verdict. When a regulator, a client or a platform asks how a clip was made, that log is your answer, and it takes thirty seconds to keep while the work is fresh.
When the voice has to change
If the local voice matters, Recast alone does not solve it. It keeps the source's soundtrack, so a different voice needs a different source recording. Shoot or generate the demo with the target-language audio first, and swap only the face. That is a production decision, not an API setting, and it is better made before the budget is spent than after the variants come back sounding the same.
Sources
Related posts
More in Use cases
- Switch voiceover language mid-video: one TTS job each, then join
ElevenLabs advertises mid-call language switching for agents. For a produced video on Sume, run one TTS job per language and join them with audio concat.
- Sync music to avatar scenes: hook, demo, CTA timestamps in the prompt
Match a Lyria bed to a 12-second multi-scene avatar video by writing [0:00-0:03] section markers into the Music Router prompt. A worked example and cost.
- Taboola Realize video ad specs: 15 s motion ad vs 90 s video
Realize (Taboola) lists a 15-second motion ad and a video spec of 6 to 30 s, 90 s max. The sizes, files and how to cut a product clip for each with Sume.
- Tag products in a YouTube Short: the sound rule and a Sume clip
Product tags in a Short: tag in the order products appear, use a Shopping sound or no sound, and block reasons like copyright claims. Plan the clip with Sume.
Written by Sume