Agent Completions: each call is a new thread, owned by your user
Every Sume Agent Completion starts a fresh thread; you cannot continue one with thread_id or send assistant turns, and runs are user-owned, not team-owned.

Each Agent Completion runs in a new thread. Sume's docs list four things as not available yet: continuing a prior thread with thread_id, assistant turns in messages[], team-owned threads, and non-image attachments. So treat each call as a one-shot task and put everything the Agent needs into that one request.
What the docs say today
These all come from the 'Not available yet' and errors sections of Agent Completions. The thread_id in the receipt identifies the new thread.
| You might try | Result today |
|---|---|
| Send an assistant turn in messages[] | The API rejects it with 400 invalid_request |
| Pass thread_id to continue a thread | Not available; each completion runs in a new thread |
| Expect team-owned threads | Completions are user-owned |
| Attach a PDF | input_image is the only attachment type |
| Stream or expect choices[] | Not available; a 202 receipt comes back |
How to design around it
Because the API rejects assistant turns instead of ignoring them, you cannot replay a past chat to give the Agent memory. Pass what it needs in instruction, or in the input field, which Sume writes to a file in the sandbox and treats as data, not instructions.
For a multi-step task, put the whole task in one completion, or run several completions and join their outputs yourself. generation_spend_cap_usd is required on every call and applies to that run only.
- Store your own conversation state in your backend.
- Send
output_schemaif you want a structured answer to feed the next call. - Poll the receipt's status; the create call returns 202, not a result.
The tradeoff
A stateless model is simpler to reason about and to cap, because each run has its own spend limit and its own sandbox. It is more work for you if you were hoping to hold a long conversation with one agent through the API. If that is the use case, Agent Completions does not fit it yet, and the docs give no date for continuation.
If you do need memory across steps, keep the state in your own store and pass a compact summary in the next call's instruction. Keep that summary free of secrets, because it goes to an agent with tool access. For one-off tasks that change each time, the docs recommend Agent Completions over a saved schedule.
Sources
Related posts
More in Developers
- Sume agent run webhook: one event per turn, not per clip
A Sume agent run sends one agent.run.terminal webhook per agent turn, never one per generated artifact. Key your handler on the run id.
- 9:16, 1080p, 8 seconds: one Sume /v1/videos request, poll and download
A copy-paste flow for a vertical clip: POST /v1/videos with aspect_ratio 9:16, resolution 1080p and duration 8, poll the job, then fetch the MP4 content URL.
- AI video audio you cannot switch off: Omni 1.1, H3 and H3 Max on Sume
Gemini Omni Flash 1.1 rejects generate_audio false; MiniMax H3 and H3 Max have no toggle. To publish a clip with your own sound, drop the audio with video trim.
- AI video generator for business: build a request form from the API
Use GET /v1/formats and the io.input_kind field to build an internal video request form, grouped by what each Sume Format needs.
Written by Sume