GPT-6 Astra async tool calls: what a slow video job means for agents
OpenAI's GPT-6 guide describes async tool calling with async: true and call_id. How that maps to a video job that takes minutes, and where Sume's job ids fit.

A video generation call is a poor fit for a tool call that must return before the model can continue, and OpenAI's GPT-6 guide adds async tool calling for cases like it: a tool call marked async: true is tracked by its call_id, so the model can keep working while the tool runs. Sume's video API is already shaped that way, because it returns a job id at once and finishes later.
The GPT-6 facts here come from OpenAI's latest-model guide and the GPT-6 Astra model page, read 2026-09-29. The guide is the only source for the async behavior; check it for the exact request fields before you build on them.
What does OpenAI say about GPT-6 and tools?
The guide says Astra's tool calling requires the Responses API even though Chat Completions is supported for other calls.
| Item | Value |
|---|---|
| Model id | gpt-6-astra |
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Price per million tokens | $10 input, $50 output |
| Tool calling API | Responses |
Why does a video job need an async pattern?
A clip takes long enough that holding one request open is fragile. Sume answers POST /v1/videos with a job id and a polling_url; hosted MCP's generate_video does the same and hands back job ids for jobs_wait.
So the tool the model sees is really two steps: start, then collect. Async tool calling lets the model treat the start as pending work instead of blocking on the result.
How do Sume's job ids line up with call_id?
They are separate ids doing separate jobs. call_id is OpenAI's handle for one tool invocation; the Sume job id names the generation. Keep a small map from one to the other in your tool code so a late result can be attached to the right call.
jobs_wait takes 1 to 20 job ids and a wait_for of all or any. A slice defaults to 50 seconds and is capped at 55. If it returns wait_slice_expired, the jobs are still running: wait again on the same ids and never call the paid create a second time.
What should I not rely on?
Do not assume a result arrives without a follow-up call. Whether your Responses code receives the collected result in the same turn or the next depends on your loop; test with a short clip first.
Keep the safety controls on regardless of the model: idempotency_key on paid calls, dry_run=true to preview, and max_spend_usd when you want a hard ceiling.
Sources
Related posts
More in Agents
- MCP tool search: how a long Sume tool list loads in Claude Code
Claude Code loads MCP tools on demand with tool search, which is on by default. What that means for Sume's long hosted tool list and how to prompt for it.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
Written by Sume