Turn a creator's raw clip into a brand-avatar ad with face swap
A workflow for re-casting your own or a licensed creator clip onto a brand avatar: trim to 4-15 s, swap on Sume, caption. Python sample and costs by tier.
To turn a raw creator clip into a brand-avatar ad on Sume, cut it to 4 to 15 seconds with clear speech, host it at a public HTTPS URL, submit it to Face Swap (Beta) with your avatar handle, then burn captions on the result. The creator's performance, timing and audio carry through; your avatar replaces the person.
Only do this with footage you own or have licensed for AI re-casting. The steps and costs come from the Face swap docs, the Video captions docs and the OpenAPI reference, read 2026-10-03.
What are the steps and what do they cost?
The avatar itself is a one-time $0.95 and must be ready before the swap. Face swap has no default quality, so pass one.
| Step | Route | Cost |
|---|---|---|
| Clear rights | Your release and licence | None |
| Cut to 4-15 s with speech | Your editor or ffmpeg | None |
| Host at a public HTTPS URL | Your storage; no signed or private URLs | None |
| Face swap | POST /v1/models/sume/avatar-face-swap/v1.0/runs | $2.76 standard, $3.68 plus, $8.25 max reserved (15 s) |
| Captions | POST /v1/video-captions | $0.20 per video up to 60 s |
| Review | Video inspect on the result | Probe and stills free |
What does the code look like?
This submits the swap and polls the job. It needs SUME_API_KEY, a ready avatar handle and a public HTTPS clip.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def main():
r = requests.post(
f"{API}/v1/models/sume/avatar-face-swap/v1.0/runs",
headers={**H, "Idempotency-Key": "creator-swap-001"},
json={
"avatar_handle": "brand_host",
"video_url": "https://example.com/creator-take-3.mp4",
"quality": "plus",
},
timeout=60,
)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(s.get("next_poll_after_seconds") or 5)
print(s["sume_status"])
if s["sume_status"] == "completed":
res = requests.get(s["result_url"], headers=H, timeout=30)
print(res.json()["data"]["linked_resource"])
main()Which clips swap well?
- One person, facing the lens, speaking clearly: the Beta expects usable speech.
- A steady camera: the workflow keeps the source's framing rather than reframing.
- A short take: about 4 to 15 seconds, so trim a long take and pick the best beat.
- Anything said is kept: the source audio is carried across, so review the words as ad copy.
How do I trim and host the clip?
Face Swap reads a public HTTPS URL; signed, expiring or private links are rejected. If the raw take is longer than 15 seconds, cut it yourself first. Sume's video trim route cuts a video that is already on media.sume.com, so a local ffmpeg cut is simpler for a clip that has not been uploaded. Then host the cut file somewhere public for the duration of the job.
If the creator posted the take publicly on TikTok or Instagram, POST /v1/media-imports can pull it onto Sume for a fixed $0.15, but only import what you have the right to re-cast. YouTube links are rejected there.
How do I budget a batch?
Face Swap reserves 15 seconds at the tier rate, so a standard swap holds $2.76 even if the clip is shorter. Ten creator takes at standard reserve $27.60 up front; add $0.20 per captioned video and a one-time $0.95 for the avatar. Use standard to pick the best takes, then re-run only the winners at plus or max.
What should I check before publishing?
- The words: the creator's speech is kept, so read the transcript as if you had written it, including any claim about the product.
- The prop: the swap keeps the product in the source, so confirm it is the right product and the right label.
- The face: look for drift across the clip with a few stills.
- The paper: confirm your licence covers re-casting, the channels and the dates you plan to run.
What then?
Run captions on the swapped video with the standalone route. A silent clip fails as caption_no_speech, but a swap keeps the source speech, so captions can read the audio. Choose a style, or leave it to the wording, and keep the clean video as your master. Then check the result with video inspect before you publish.
Label the ad as AI-generated where the platform asks, and keep the release, the original take and the job id on file. If you want the avatar to say words the creator did not say, a swap is the wrong tool; use a talking video and write the script yourself.
Sources
Related posts
- Face swap a video with a photo: create an avatar first
- Avatar video captions: inline add-on or standalone $0.20 job?
- Review a creator clip's transcript before reusing it as a paid ad
- TikTok impersonation deepfake ban: use face swap with your own avatar
- AI face swap video cost: the $2.76 to $8.25 ceiling per clip
More in Use cases
- Does a voiceover make a reused YouTube Short original?
No: YouTube's Oct 1, 2026 update says narration describing what's on screen doesn't make a re-upload original. What to change instead, and how Sume helps.
- Does adding music make a YouTube Short original? Oct 2026 rule
YouTube now cuts reach for re-uploaded Shorts. A new soundtrack on borrowed footage is a minor edit; here is what counts as original and how Sume fits.
- Does flipping or cropping a video make it original on Shorts?
No. YouTube's Oct 1, 2026 update names minor technical edits as not enough. What Sume's crop and flip tools are for, and what makes a Short yours.
- Does TikTok keep Content Credentials on a downloaded AI video?
TikTok says it is working to attach Content Credentials that stay on download. What TikTok says, what its 2024 timeline promised, and how to test a file.
Written by Sume