Porting chat-completions code to Sume Agent Completions

Agent Completions takes system and user messages, rejects assistant turns, returns a 202 receipt, and does not stream. Here is what to change when porting.

4 min readSume
All posts

Porting chat-completions code to Sume Agent Completions means four changes: send only system and user messages, expect a 202 receipt instead of a reply, poll or use a webhook instead of streaming, and set a required spend cap. Assistant turns are rejected and there is no continuation by thread_id.

The endpoint is POST /v1/agent/completions. It starts a video agent run and returns an agrun_ receipt with object agent.run and model sume-agent.

Most porting pain comes from the first and last habit: expecting an answer in the response and forgetting that a paid run needs a budget. Everything else follows from those two.

What stays the same

You send exactly one of instruction or messages. With messages, the shape is familiar: system and user turns. You can add input data, up to 30 image attachments, an output_schema with a primary_output_key, and a communication.webhook_url. An Idempotency-Key works as on Format runs. See Agent Completions.

If your existing code builds a messages array with prior assistant replies, drop those turns and fold anything the run needs into the system message or the input object. An assistant turn in the array is rejected rather than ignored, so you will find the mistake on the first call.

What changes

There is no synchronous choices array. The response is a 202 receipt and you read the finished run later from GET /v1/agent-runs/{id} or the status route, or cancel at /v1/agent-runs/{id}/cancel. GET /v1/agent-runs lists runs.

Each completion runs in a new thread, so a follow-up is a fresh request, not an assistant message appended to a history. Attachments are images only. Completions are user-owned, with no team threads, and service-account keys are refused. Scopes are agent_completions:read and agent_completions:write.

Because nothing streams, design the interface around waiting. Show a status line, poll with a backoff that doubles up to 60 seconds, and give the user a way to cancel. If your product is a chat box, tell people up front that a video takes minutes, not seconds.

Chat completions habits versus Agent Completions, Sume docs read 2026-10-10
HabitAgent Completions
Reply in the response body202 receipt, read the run later
stream: trueNo streaming
Assistant turns in messagesRejected
Continue by conversation idNo, each completion is a new thread
Optional token limitgeneration_spend_cap_usd is required
Non-image filesNot supported as attachments

The required cap

generation_spend_cap_usd has no default on this endpoint. Leaving it out returns 400 invalid_request. Set it from the first call. The reason is simple: an agent run can call paid media tools, so the budget travels with the request.

Another trap is the id family. Passing an Action or Format run id to the agent-run routes gives 404 agent_run_not_found, because each family has its own routes.

Use a webhook for production and polling for development. A webhook arrives when the run reaches a terminal state, so the code that handles it is small: verify, dedupe on request_id, and store the result.

A minimal request

The call below starts a run with a one dollar cap. The webhook is optional; omit it and poll instead.

import json, os, urllib.request

body = {
    "instruction": "Make a 15 second teaser from the attached product photo.",
    "generation_spend_cap_usd": 1.0,
}
req = urllib.request.Request(
    "https://api.sume.com/v1/agent/completions",
    data=json.dumps(body).encode(),
    headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
             "Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=30) as resp:
    print(resp.status, json.load(resp).get("id"))

Sources

Related posts

More in Agents

All Agents posts

Written by Sume