Porting chat history to Agent Completions: assistant turns fail

Sume Agent Completions take OpenAI-style messages but return 400 on any assistant turn and start each call in a new thread. What to send instead.

5 min readSume
All posts

If you send an OpenAI-style conversation to POST /v1/agent/completions with a prior assistant message in the array, the API rejects it with 400 invalid_request. It does not ignore the turn. The docs give the reason: accepting assistant turns would imply that Sume replays a prior conversation, and this endpoint does not do that today.

Each completion also runs in a new thread. thread_id appears in the receipt, but you cannot pass it back to continue. So a chat client has to collapse its history into the current task.

What to send

Use system and user turns only, or a single instruction string. Send exactly one of instruction or messages. Sume joins the turns in order into one prompt. content may be a string, an OpenAI-style [{type: "text"}] array, or input_text and input_image parts.

Put prior context into the user message as text you control, and put bulky data in input, which Sume writes to a workspace file and treats as data, not instructions. That keeps a pasted transcript from being read as new orders.

Messages handling, as of 2026-10-09 (docs.sume.com/agents/completions)
You sendResult
system and user turnsJoined in order into one prompt
An assistant turn400 invalid_request
Both instruction and messages400 invalid_request
thread_idNot supported; each call starts a new thread
input_image parts or attachmentsUp to 30 images the agent can see
stream: true or choices[] expectationNot available; response is an agent.run receipt

A rewrite pattern

Summarize the earlier turns in your own code, then send one user message that states the current ask and the relevant facts. For tasks that must remember state across calls, store the state yourself and pass it in input, or use a Format with previous_run_id where that fits, since continued Format runs are a separate documented path.

  • Strip assistant turns before the call; do not rely on the API to ignore them.
  • Set generation_spend_cap_usd; it is required and has no default.
  • Poll status_url or register a webhook; do not wait for a synchronous answer.
  • Use output_schema when your code needs fields, not prose.

Why this is not a chat drop-in

The docs say it plainly: this is not a synchronous chat completion. A real agent turn opens a sandbox, calls tools and may generate media, which takes too long for one HTTP request. Plan the integration as a job queue that happens to accept familiar request syntax.

Checklist for the port

List what your chat client sends today: roles, tool messages, images, system prompt, streaming flag. For each, the table above says whether Agent Completions accept it. The usual changes are removing assistant turns, turning streaming into polling or a webhook, and replacing a thread id with your own stored state.

Test with a tiny cap first. A request that includes an unknown top-level field on the Format surface returns unknown_parameter; on this endpoint the documented 400 causes are a missing cap, both or neither of instruction and messages, an assistant turn, a malformed input, or a model other than sume-agent.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume