Agent Completions gaps today: plan around streaming and PDFs
What Sume Agent Completions does not do yet: streaming, thread continuation, assistant turns, PDFs. Workarounds for GPT-6.1 Sol and Sonnet 5.5 agents.

Sume's Agent Completions page lists what is not available yet: non-image attachments (images are the only supported type), streaming and a synchronous OpenAI-style choices[] response, continuing a prior thread with thread_id and assistant turns, and team-owned threads. If your GPT-6.1 Sol or Sonnet 5.5 agent expects any of those, design around them now instead of finding out in production.
The list and the workaround
This is taken from the "Not available yet" section of the Agent Completions docs. Workarounds below are design choices, not features; nothing in them depends on unreleased behavior.
| Not available yet | Plan around it by | Cost of the workaround |
|---|---|---|
| Non-image attachments such as PDFs | Putting the text in input as data | You extract the text yourself |
| Streaming | Polling status_url, or a run webhook | Latency of the poll interval |
choices[] response | Reading output.text from the receipt | Not a drop-in for chat clients |
| Continuing a thread | Starting a new run with a summary | The agent has no memory of earlier runs |
assistant turns | Flattening history into one user turn | Longer prompts |
| Team-owned threads | Running under a user-owned key | Runs belong to a person |
Completion is pushed, progress is not
A finished run can be pushed. A signed POST goes to your public HTTPS communication.webhook_url with up to ten attempts and a ten second timeout, as described in Run webhooks. Progress is not pushed, so a long task needs polling in the meantime. Keep the webhook as the fast path and a poll as the backstop.
Design rules
- Make each run self-contained: the instruction, the data in
input, the cap and the schema. Assume no memory. - Persist your own state between runs, keyed by your ids, and pass only what is needed.
- Bind
output_schemawhen a program will read the result, per Structured output. - Keep the cap per run; a pipeline of runs gets a cap on each, plus a total you track yourself.
- Read finished media from the durable URLs in
outputrather than rendering inline.
When to revisit
Check the page again before building anything that assumes a gap closed; this list is the contract, and it can change. For work that needs conversation, a Format or an interactive chat in the app may fit better, and Jobs and results shows how to read what runs produced. Keep API keys and signed URLs out of logs, per Safe automation.
Sources
Related posts
More in Agents
- AI video agent vs a single-model video generator: what you call
A single-model generator returns one clip from a prompt. An agent plans shots, calls tools and assembles a video. How the Sume calls differ.
- @-mention an agent on a video asset: Runway vs Sume
Runway Enterprise lets you @-mention its Agent in asset comments. Sume Agents take work via the Agent Completions API; media jobs report by webhook.
- Re-render Sora prompts from an agent: jobs_wait takes 20 ids
An agent re-rendering saved Sora prompts should wait on up to 20 Sume job ids per call, read results in one batch, and not resubmit after a wait slice expires.
- Claude Code 2.1.289 agent.spawn: teammates share one Sume queue
Claude Code 2.1.289 adds agent.spawn for teammates. Teammates sharing a Sume workspace share its concurrency limit and queue, so plan the width of the fan-out.
Written by Sume