GPT-6 Astra async tool calls: what a slow video job means for agents

OpenAI's GPT-6 guide describes async tool calling with async: true and call_id. How that maps to a video job that takes minutes, and where Sume's job ids fit.

4 min readSume
All posts

A video generation call is a poor fit for a tool call that must return before the model can continue, and OpenAI's GPT-6 guide adds async tool calling for cases like it: a tool call marked async: true is tracked by its call_id, so the model can keep working while the tool runs. Sume's video API is already shaped that way, because it returns a job id at once and finishes later.

The GPT-6 facts here come from OpenAI's latest-model guide and the GPT-6 Astra model page, read 2026-09-29. The guide is the only source for the async behavior; check it for the exact request fields before you build on them.

What does OpenAI say about GPT-6 and tools?

The guide says Astra's tool calling requires the Responses API even though Chat Completions is supported for other calls.

From OpenAI's GPT-6 Astra model page and latest-model guide, read 2026-09-29.
ItemValue
Model idgpt-6-astra
Context window1,050,000 tokens
Max output128,000 tokens
Price per million tokens$10 input, $50 output
Tool calling APIResponses

Why does a video job need an async pattern?

A clip takes long enough that holding one request open is fragile. Sume answers POST /v1/videos with a job id and a polling_url; hosted MCP's generate_video does the same and hands back job ids for jobs_wait.

So the tool the model sees is really two steps: start, then collect. Async tool calling lets the model treat the start as pending work instead of blocking on the result.

How do Sume's job ids line up with call_id?

They are separate ids doing separate jobs. call_id is OpenAI's handle for one tool invocation; the Sume job id names the generation. Keep a small map from one to the other in your tool code so a late result can be attached to the right call.

jobs_wait takes 1 to 20 job ids and a wait_for of all or any. A slice defaults to 50 seconds and is capped at 55. If it returns wait_slice_expired, the jobs are still running: wait again on the same ids and never call the paid create a second time.

What should I not rely on?

Do not assume a result arrives without a follow-up call. Whether your Responses code receives the collected result in the same turn or the next depends on your loop; test with a short clip first.

Keep the safety controls on regardless of the model: idempotency_key on paid calls, dry_run=true to preview, and max_spend_usd when you want a hard ceiling.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume