Agent Completions rejects assistant turns: how to carry chat history
Sume Agent Completions return 400 on any assistant message in messages[] and start a fresh thread each time. Here is what that means for multi-turn history.

If you send an assistant turn in messages[] to Sume Agent Completions, the request fails with 400 invalid_request: the API rejects such turns, it does not ignore them. The reason in the docs is that accepting them would imply Sume replays a prior conversation, and this endpoint does not do that. Each completion runs in a new thread.
So an OpenAI-style client that replays its whole chat on every call will fail on the first call after a model reply. The fix is to shape the history yourself before you send it.
What the endpoint accepts
From the Agent Completions page, read 2026-10-09.
| Request element | Accepted? | Notes |
|---|---|---|
system turn | Yes | Joined in order into one prompt |
user turn | Yes | String or [{type:"text"}]; input_text and input_image parts also work |
assistant turn | No | 400 invalid_request |
thread_id continuation | Not available | Listed under not available yet |
Both instruction and messages | No | Send exactly one |
generation_spend_cap_usd | Required | No default |
Carrying context without assistant turns
Three workable patterns follow from the documented limits. These are suggestions, not documented features.
First, fold the earlier exchange into one user message as quoted text: what was asked, what came back, what to do next. Second, put structured prior results in the input object, which Sume writes to /workspace/inputs/sume-action-input.json and treats as data, not instructions. Third, pass durable results, such as a previous run's media URLs, as text in the instruction or as image attachments.
Cost and safety notes
Every call needs its own spend cap, so a long conversation is a sequence of separately capped runs; add the caps up to know the worst case. Keep untrusted pasted text in input rather than in the prompt where possible, since the docs say Sume uses input only as data.
Use an Idempotency-Key on each call. Sending the same key again with the same payload returns the original receipt with idempotency_hit: true; a different payload under the same key is 409 idempotency_conflict.
A small illustration
Suppose a chat UI has three exchanges and the user now asks for a darker version of the last image. Do not send the earlier replies as assistant turns. Send one user message that states the original brief, names the image URL that came back, and gives the new instruction, with a spend cap sized for one image. If the earlier image is the thing being edited, attach it as an input_image part so the agent can actually look at it.
This keeps the request inside what the endpoint accepts and makes the context explicit, which also makes the run easier to audit later from the agent.run receipt. Keep the folded history short. Long pasted transcripts raise token use inside the agent turn and can bury the new instruction, so summarize earlier turns in a few lines and link or attach only the artifacts the next step needs.
Sources
Related posts
More in Agents
- Agent run log allowlist: which Sume fields are safe to keep
Log request ids, job ids, status and sanitized media metadata from Sume agent runs; never log API keys, signed URLs, raw private media URLs or full transcripts.
- Should an agent tool take a workspace_id argument for Sume calls?
No. Sume's safe-automation guidance says the API key or app session selects the workspace and tools must not take a workspace id from the user.
- Animate a product photo with Agent Completions: $0.625 on Wan 3.0
Send one input_image and ask the Sume agent for a 5-second Wan 3.0 clip at 720p: $0.625 at $0.125 per second. Cap 1 is enough. Body and limits.
- Sume API key with actions:read only: list schedules, not start runs
A Sume key with actions:read can list and read schedules and runs. Starting or canceling a run needs actions:write; keys made before the trigger lack both.
Written by Sume