Voice agent hand-off: ask for a clip, get an async Sume job
A live voice agent should not wait on a video render. Hand the request to an async Sume job, speak the job id back, and deliver the clip by poll or webhook.

When a voice agent is asked for a video clip mid-conversation, submit a Sume job in async mode, keep the job id, tell the caller the clip is on its way, and deliver the result later by webhook or status polling. Do not hold the live audio session open while the render runs.
Why live voice and video jobs do not share a clock
The OpenAI API changelog for Sep 10, 2026 says GPT-Live 1 became generally available in the API for full-duplex voice conversations. Full-duplex means the model listens and speaks at the same time, so the session is judged in conversational turns.
Sume generation is the opposite shape. Every submit endpoint creates a durable job, and the docs state that video, avatar-video and face-swap jobs routinely outlast the 30 second wait budget of sync mode. A 2xx response only means the job exists and paid work is in flight.
The hand-off, step by step
The voice model calls one tool. The tool handler does the Sume work and returns immediately with a short status the model can speak.
- Submit to
POST /v1/videoswithmodel: "sume/auto"and anIdempotency-Key, so a retried tool call returns the original job instead of billing a second one. - Pass
callback_url(HTTPS) if your server can receive the terminal webhook, and keep pollingGET /v1/jobs/{id}/statusas a backup. - Return the job id to the voice model, which tells the caller the clip is rendering.
- On
completed, read the result and send the media link through a channel the caller can open, such as a message or email.
A tool handler that returns at once
This Python handler submits the job and returns the id. It uses only documented fields from the Video Generation page. Auto accepts 3 to 10 second clips at 16:9 or 9:16.
import os
import uuid
import requests
API = "https://api.sume.com"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def request_clip(prompt: str, callback_url: str | None = None) -> dict:
body = {
"model": "sume/auto",
"prompt": prompt,
"aspect_ratio": "9:16",
"duration": 5,
}
if callback_url:
body["callback_url"] = callback_url
headers = {**HEADERS, "Idempotency-Key": f"voice-{uuid.uuid4()}"}
r = requests.post(f"{API}/v1/videos", json=body, headers=headers, timeout=30)
r.raise_for_status()
job = r.json()
return {"job_id": job["id"], "say": "Your clip is rendering. I will send it when it is ready."}
def job_status(job_id: str) -> dict:
r = requests.get(f"{API}/v1/jobs/{job_id}/status", headers=HEADERS, timeout=30)
r.raise_for_status()
return r.json()Facts used in this post
Both rows below were read on 2026-10-03.
| Surface | Shape | What your tool does |
|---|---|---|
| GPT-Live 1 (OpenAI, GA Sep 10, 2026) | Full-duplex voice conversation | Calls your tool and keeps talking |
Sume async job | Durable job, poll or webhook | Returns the job id at once |
Sume sync mode | Waits at most 30 seconds | Not suited to video |
A client-side timeout never cancels a Sume job; it keeps running and keeps billing. Store the id so a restarted process can pick the job up again. See Jobs and results for the full status list.
Sources
Related posts
More in Developers
- VS Code 1.140 shared MCP config files: what goes in the Sume entry
VS Code 1.140 lets MCP servers live in portable config files shared across Copilot tools. For Sume the entry is one URL, and no key belongs in the file.
- waitForJob throws on a failed poll: resume by job id in TypeScript
Unlike waitForRun, waitForJob has no transient-failure budget: one failed status read after the client's retries throws. Wrap it and resume by job id.
- waitForRun maxTransientFailures: how many bad polls it absorbs
waitForRun absorbs 6 consecutive 429, 5xx or network read failures before it throws. How the streak resets, backoff works, and onTransientError fits in.
- Wan 3.0 at 30 fps and 2 to 30 s: the Sume model id
Model Studio lists Wan 3.0 at 480P to 1080P, 2 to 30 s, 30 fps. On Sume the id is wan-3.0, 2 to 30 s; first and last frames go in frame_images.
Written by Sume