Agent Completions takes messages[] but returns a receipt

Sume Agent Completions borrows the OpenAI messages[] request shape, yet create returns 202 and a run receipt to poll: no choices[] and no streaming.

5 min readSume
All posts

Sume Agent Completions accepts messages[] in an OpenAI-like shape, but it is not a drop-in chat completions endpoint. The create call returns 202 with a run receipt, you poll the run, and the final answer is the run's output, not a choices[] array. Streaming and a synchronous OpenAI-compatible response are listed as not available yet.

The details are from Agent Completions, read on 2026-10-03.

Why it is asynchronous

A real agent turn opens a sandbox, calls tools and may generate media. That takes longer than an HTTP request should stay open, so Sume accepts the work and hands back a receipt. The request borrows the messages[] shape so existing plumbing, such as a message-building helper, can be reused, but the response side is a run receipt.

That has a direct consequence for code ported from a chat SDK. A client that reads response.choices[0].message.content will find nothing. The port needs a poll loop and a place to put the result.

Chat completions habit versus Agent Completions (read 2026-10-03)
Habit from chat APIsAgent Completions
Synchronous response202 receipt, then poll
choices[] in the bodyRun receipt; output comes from the run
Stream tokensStreaming not available yet
assistant turns in messages[]Rejected; assistant turns are not available yet
Any model idmodel must be sume-agent
Spend is implicitgeneration_spend_cap_usd is required

What a port needs

Three changes cover most ports. First, send generation_spend_cap_usd on every request; a missing cap is a 400 invalid_request. Second, send an Idempotency-Key so a retried create does not start a second run; reusing a key with a different payload is a 409 idempotency_conflict. Third, poll the run until it is terminal, using an agent_completions:read key.

The scopes are separate from the rest of the API: agent_completions:write creates and cancels, agent_completions:read polls. A key without them gets 403 insufficient_scope, and so does a service-account key.

What it still does that chat does not

The run can attach images as input_image content parts and can parse the result against an output_schema you supply, so the output is structured data rather than prose to be scraped. Those two features compose: the images reach the agent, and the output is still parsed against your schema after the run completes.

If your use case really is a single synchronous model call with streaming, this endpoint is the wrong tool today. If it is an agent that can look at an image, run tools and produce media under a spend ceiling, the asynchronous shape is the point.

What you cannot do yet

The docs list four gaps. Only images can be attached: input_image is the only content type today, and PDFs and other files come later. There is no streaming and no synchronous choices[] response. You cannot continue a prior thread with thread_id and you cannot send assistant turns in messages[]. And completions are user-owned, so there are no team-owned threads.

Plan around these before porting. A chat UI that replays assistant history into each request will be rejected, so keep conversation state in your own app and send a single fresh instruction with the context it needs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume