AI presenter video from a script: first clip in Python on Sume

Make an AI presenter clip from a script in one Python file: submit to /v1/avatar-1.0/talking-video, poll to terminal, read the result. Costs and limits.

5 min readSume
All posts

To make an AI presenter video from a script, call POST /v1/avatar-1.0/talking-video with an avatar_handle and a script, then poll the job until it is terminal and read the result. The Python file below does that with only the standard library. A 15-second Standard clip costs about $2.76 at the listed rate (read 2026-10-03), and the avatar handle must already exist.

Script-to-video is the whole product surface for most presenter use cases: onboarding, announcements, release notes, FAQ answers. The workflow is asynchronous, so the client has to poll, and the one thing to get right is idempotency.

The request

Avatar creation and video generation are two steps. First create the avatar at POST /v1/avatar-1.0/generate with a handle and an input of type prompt, props or photo, then wait for that job. Once the avatar is ready, the talking video route takes avatar_handle, exactly one of script or video_inputs, and optional quality, aspect_ratio, scene, product_image and captions.

The script must estimate at 4 to 60 seconds. quality defaults to plus if you omit it, so the example sets standard explicitly. Always send an Idempotency-Key on submit: if the connection drops, retrying with the same key returns the original job rather than billing a second one.

import json, os, sys, time, urllib.request

API = "https://api.sume.com"
KEY = os.environ.get("SUME_API_KEY", "")
if not KEY:
    sys.exit("set SUME_API_KEY first")

def call(method, path, body=None, key=None):
    req = urllib.request.Request(API + path, method=method,
        data=json.dumps(body).encode() if body else None,
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json",
                 **({"Idempotency-Key": key} if key else {})})
    with urllib.request.urlopen(req) as r:
        return json.load(r)

job = call("POST", "/v1/avatar-1.0/talking-video", {
    "avatar_handle": "product_host",
    "quality": "standard",
    "script": "Meet the new launch checklist: three steps, one page, ready to share.",
}, key="presenter-first-clip-001")
job_id = job["request_id"]
while True:
    s = call("GET", f"/v1/jobs/{job_id}/status")
    if s.get("terminal"):
        break
    time.sleep(s.get("next_poll_after_seconds") or 5)
print(s.get("sume_status"), call("GET", f"/v1/jobs/{job_id}/result") if s.get("sume_status") == "completed" else "")

The polling loop

The documented client pattern is: submit with the default async mode, read the job id from the envelope, then GET /v1/jobs/{id}/status until terminal is true, waiting for next_poll_after_seconds when present. A completed job's result carries public media.sume.com video artifacts and preview fields; failed jobs expose public error metadata. Read the full flow in Jobs and results.

Do not submit a new paid job for the same intent while waiting. If the client crashes, rerun with the same idempotency key and the original job comes back.

What to add before production

Keep the script short at first. Fifteen seconds is long enough to judge the avatar and cheap enough to retake.

  • A balance check: read GET /v1/balance before a batch so a 402 insufficient_credits does not surprise you.
  • Webhooks: use the documented webhook mode instead of polling if you run many jobs.
  • Captions: add captions to burn styled text into the final MP4; failure there does not fail the whole job.
  • Aspect ratio: the default is 9:16; set 16:9 for slides and webinars.

Common first-run mistakes

The most frequent failure is sending both script and video_inputs: the route requires exactly one. The second is a script that estimates over 60 seconds, which is rejected; split it into several jobs. Third is an avatar that is not ready yet, because the avatar job has not completed: poll the creation job first and use the handle only afterwards.

Finally, remember the handle rules: it may include a leading @, but Sume stores it without one, so use the same normalised spelling in your own records.

Cost of the first run

A cautious first run is one avatar and one 15-second Standard clip: $0.95 plus $2.76, about $3.71 in total at listed rates (read 2026-10-03). Plus costs $3.675 for the same clip and Max $8.25. If the first result is close but not right, change the script and resubmit; the avatar does not need to be recreated, because the handle is reusable.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume