AI presenter video from a script: first clip in Python on Sume
Make an AI presenter clip from a script in one Python file: submit to /v1/avatar-1.0/talking-video, poll to terminal, read the result. Costs and limits.

To make an AI presenter video from a script, call POST /v1/avatar-1.0/talking-video with an avatar_handle and a script, then poll the job until it is terminal and read the result. The Python file below does that with only the standard library. A 15-second Standard clip costs about $2.76 at the listed rate (read 2026-10-03), and the avatar handle must already exist.
Script-to-video is the whole product surface for most presenter use cases: onboarding, announcements, release notes, FAQ answers. The workflow is asynchronous, so the client has to poll, and the one thing to get right is idempotency.
The request
Avatar creation and video generation are two steps. First create the avatar at POST /v1/avatar-1.0/generate with a handle and an input of type prompt, props or photo, then wait for that job. Once the avatar is ready, the talking video route takes avatar_handle, exactly one of script or video_inputs, and optional quality, aspect_ratio, scene, product_image and captions.
The script must estimate at 4 to 60 seconds. quality defaults to plus if you omit it, so the example sets standard explicitly. Always send an Idempotency-Key on submit: if the connection drops, retrying with the same key returns the original job rather than billing a second one.
import json, os, sys, time, urllib.request
API = "https://api.sume.com"
KEY = os.environ.get("SUME_API_KEY", "")
if not KEY:
sys.exit("set SUME_API_KEY first")
def call(method, path, body=None, key=None):
req = urllib.request.Request(API + path, method=method,
data=json.dumps(body).encode() if body else None,
headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json",
**({"Idempotency-Key": key} if key else {})})
with urllib.request.urlopen(req) as r:
return json.load(r)
job = call("POST", "/v1/avatar-1.0/talking-video", {
"avatar_handle": "product_host",
"quality": "standard",
"script": "Meet the new launch checklist: three steps, one page, ready to share.",
}, key="presenter-first-clip-001")
job_id = job["request_id"]
while True:
s = call("GET", f"/v1/jobs/{job_id}/status")
if s.get("terminal"):
break
time.sleep(s.get("next_poll_after_seconds") or 5)
print(s.get("sume_status"), call("GET", f"/v1/jobs/{job_id}/result") if s.get("sume_status") == "completed" else "")
The polling loop
The documented client pattern is: submit with the default async mode, read the job id from the envelope, then GET /v1/jobs/{id}/status until terminal is true, waiting for next_poll_after_seconds when present. A completed job's result carries public media.sume.com video artifacts and preview fields; failed jobs expose public error metadata. Read the full flow in Jobs and results.
Do not submit a new paid job for the same intent while waiting. If the client crashes, rerun with the same idempotency key and the original job comes back.
What to add before production
Keep the script short at first. Fifteen seconds is long enough to judge the avatar and cheap enough to retake.
- A balance check: read
GET /v1/balancebefore a batch so a402 insufficient_creditsdoes not surprise you. - Webhooks: use the documented webhook mode instead of polling if you run many jobs.
- Captions: add
captionsto burn styled text into the final MP4; failure there does not fail the whole job. - Aspect ratio: the default is
9:16; set16:9for slides and webinars.
Common first-run mistakes
The most frequent failure is sending both script and video_inputs: the route requires exactly one. The second is a script that estimates over 60 seconds, which is rejected; split it into several jobs. Third is an avatar that is not ready yet, because the avatar job has not completed: poll the creation job first and use the handle only afterwards.
Finally, remember the handle rules: it may include a leading @, but Sume stores it without one, so use the same normalised spelling in your own records.
Cost of the first run
A cautious first run is one avatar and one 15-second Standard clip: $0.95 plus $2.76, about $3.71 in total at listed rates (read 2026-10-03). Plus costs $3.675 for the same clip and Max $8.25. If the first result is close but not right, change the script and resubmit; the avatar does not need to be recreated, because the handle is reusable.
Sources
Related posts
More in Developers
- AI SDK MCP tools(): explicit Zod schemas for Sume's jobs_wait
Pass explicit Zod schemas to mcpClient.tools() so your app exposes only the Sume tools it needs, with typed job ids and a bounded wait_for.
- AI video defaults on Sume: 768p, 480p, 720p and 8 seconds explained
Recast defaults to 768p, Genjutsu to 480p, Gemini Omni Flash and sume/auto to 720p and 8 s. Which defaults change price, and the one place fal differs.
- AI video models with no aspect_ratio option: Recast, Genjutsu, Grok
Three Sume video rows reject aspect_ratio: H3 Max Recast, Higgsfield Genjutsu and Grok Imagine Video 1.5. What sets the output frame instead, and what to send.
- AI video person swap rejects long takes: find shots over 15 s
H3 Max Recast on Sume needs 5 to 30 seconds with no single shot over 15. Find the long takes with ffmpeg scene detection, then cut them before you pay.
Written by Sume