English avatar video with Spanish and Korean subtitles via cues
Make the avatar clip in English, then burn Spanish and Korean subtitles with video-captions cues. Korean needs a Hangul style such as black-outline, not slam.
Sume Avatar 1.0 speaks English, but the picture can carry other languages as subtitles. Generate the clip once, then call POST /v1/video-captions once per language with authored cues, using black-outline for Korean because the Latin styles reject Hangul.
Why cues and not speech-to-text
If you let a caption job transcribe the soundtrack, you get English. Passing cues (each with text, start, end in seconds) skips speech-to-text and burns exactly your copy, so the translated text is whatever you supply. cues, segments, words and script_text are mutually exclusive.
| Language | Style to name | Why |
|---|---|---|
| Spanish | slam | Latin glyphs, default look |
| Korean | black-outline | Hangul face; the Latin styles return 400 caption_hangul_text_latin_style |
One job per language
Each job is billed separately at $0.20 for a clip up to 60 seconds, so two languages is $0.40 on top of the avatar clip. Use a distinct Idempotency-Key per language so a retry cannot double-bill.
Set VIDEO_URL to the public HTTPS URL of the finished avatar video (a caption job needs a fetchable URL).
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
VIDEO = os.environ["VIDEO_URL"]
JOBS = {
"es": ("slam", [{"text": "Hola, soy tu guía.", "start": 0.0, "end": 2.2}]),
"ko": ("black-outline", [{"text": "안녕하세요, 안내를 맡았습니다.", "start": 0.0, "end": 2.2}]),
}
for lang, (style, cues) in JOBS.items():
r = requests.post(
"https://api.sume.com/v1/video-captions",
headers={**H, "Idempotency-Key": f"avatar-subs-{lang}-001"},
json={"video_url": VIDEO, "style": style, "cues": cues},
timeout=30,
)
r.raise_for_status()
print(lang, r.json()["status_url"])Keep the timings honest
The spoken English sets the clock. Match each translated cue to the English phrase it translates, and read the line aloud at that length; Korean and Spanish often run longer or shorter than the English they replace.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar face swap Beta: check the 4 to 15 second clip before you submit
Sume's Avatar Face Swap Beta needs a public HTTPS source clip of about 4 to 15 seconds with usable audio and a required quality tier.
- Does Sume face swap keep the original voice, outfit and scene?
Sume's Face Swap (Beta) keeps the source clip's audio, camera and timing but replaces the whole person, not just the face. What stays, what changes, cost.
- Face swap API webhook: get job.completed instead of polling
Sume's face swap (Beta) takes mode webhook with a public HTTPS webhook_url, then sends one signed terminal event: job.completed, job.failed or job.canceled.
- Face swap for UGC ads: source clip rules (Sume beta)
Sume's beta Avatar Face Swap applies a ready avatar face to a public 4-15 second source video with audio. Required fields, quality tiers and what it rejects.
Written by Sume