Voice agent hand-off: ask for a clip, get an async Sume job

A live voice agent should not wait on a video render. Hand the request to an async Sume job, speak the job id back, and deliver the clip by poll or webhook.

4 min readSume
All posts

When a voice agent is asked for a video clip mid-conversation, submit a Sume job in async mode, keep the job id, tell the caller the clip is on its way, and deliver the result later by webhook or status polling. Do not hold the live audio session open while the render runs.

Why live voice and video jobs do not share a clock

The OpenAI API changelog for Sep 10, 2026 says GPT-Live 1 became generally available in the API for full-duplex voice conversations. Full-duplex means the model listens and speaks at the same time, so the session is judged in conversational turns.

Sume generation is the opposite shape. Every submit endpoint creates a durable job, and the docs state that video, avatar-video and face-swap jobs routinely outlast the 30 second wait budget of sync mode. A 2xx response only means the job exists and paid work is in flight.

The hand-off, step by step

The voice model calls one tool. The tool handler does the Sume work and returns immediately with a short status the model can speak.

  • Submit to POST /v1/videos with model: "sume/auto" and an Idempotency-Key, so a retried tool call returns the original job instead of billing a second one.
  • Pass callback_url (HTTPS) if your server can receive the terminal webhook, and keep polling GET /v1/jobs/{id}/status as a backup.
  • Return the job id to the voice model, which tells the caller the clip is rendering.
  • On completed, read the result and send the media link through a channel the caller can open, such as a message or email.

A tool handler that returns at once

This Python handler submits the job and returns the id. It uses only documented fields from the Video Generation page. Auto accepts 3 to 10 second clips at 16:9 or 9:16.

import os
import uuid
import requests

API = "https://api.sume.com"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}


def request_clip(prompt: str, callback_url: str | None = None) -> dict:
    body = {
        "model": "sume/auto",
        "prompt": prompt,
        "aspect_ratio": "9:16",
        "duration": 5,
    }
    if callback_url:
        body["callback_url"] = callback_url
    headers = {**HEADERS, "Idempotency-Key": f"voice-{uuid.uuid4()}"}
    r = requests.post(f"{API}/v1/videos", json=body, headers=headers, timeout=30)
    r.raise_for_status()
    job = r.json()
    return {"job_id": job["id"], "say": "Your clip is rendering. I will send it when it is ready."}


def job_status(job_id: str) -> dict:
    r = requests.get(f"{API}/v1/jobs/{job_id}/status", headers=HEADERS, timeout=30)
    r.raise_for_status()
    return r.json()

Facts used in this post

Both rows below were read on 2026-10-03.

Realtime voice versus Sume jobs (read 2026-10-03)
SurfaceShapeWhat your tool does
GPT-Live 1 (OpenAI, GA Sep 10, 2026)Full-duplex voice conversationCalls your tool and keeps talking
Sume async jobDurable job, poll or webhookReturns the job id at once
Sume sync modeWaits at most 30 secondsNot suited to video

A client-side timeout never cancels a Sume job; it keeps running and keeps billing. Store the id so a restarted process can pick the job up again. See Jobs and results for the full status list.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume