Agent Completion or three API calls for a render-trim-caption chain
If the steps are fixed, call the endpoints. If the task changes on every call, an Agent Completion with a required spend cap fits. How the two compare.

Use three API calls when the pipeline is fixed: generate, trim, then captions, each with its own key. Use an Agent Completion when the task changes on every call and you want the agent to decide the steps. The Agent Completions page says it directly: a Format stores how to do something, a schedule stores what to do, and an Agent Completion stores nothing, so you send the task each time.
What the two approaches give you
An Agent Completion runs the same agent as the Agents chat UI, with a sandbox, tools and media generation. It is asynchronous: POST /v1/agent/completions returns 202 and a receipt with status_url and cancel_url, and you poll GET /v1/agent-runs/$RUN_ID. Streaming and a sync OpenAI-compatible wire are not available.
| Question | Three endpoint calls | Agent Completion |
|---|---|---|
| Who decides the steps | Your code | The agent |
| What you poll | GET /v1/jobs/:id/status per step | GET /v1/agent-runs/$RUN_ID |
| Spend control | You pick each step and its price | generation_spend_cap_usd, required, no default |
| Webhook event | job.completed per job | One agent.run.terminal per run |
| Idempotency | Idempotency-Key per submit | Idempotency-Key on the create; a replay returns the original receipt |
| Stored by Sume | Jobs and artifacts | Nothing about the task itself |
The spend cap replaces the approval prompt
generation_spend_cap_usd has no default, and a create without it fails with 400 invalid_request. The docs explain why: an Agent Completion is an unattended agent with access to your generation wallet. In the chat UI an approval prompt protects you, and the cap is what replaces that prompt for a backend caller. Set it to the most you are willing to spend on that single run.
A fixed chain has a natural ceiling because you know the steps: one render, one trim at $0.02, one caption at $0.20. An agent can take more steps than you planned, and the cap is how you bound that.
A request, from the docs
This is the documented create. Poll status_url until the value of next_action is no longer poll_status.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"instruction": "Make two product stills from https://example.com/p and summarize what you shot.",
"generation_spend_cap_usd": 5
}'Webhooks and cancel
A run delivers exactly one terminal webhook, agent.run.terminal, with outcome of ok, degraded or error. A canceled run delivers no webhook, so after POST /v1/agent-runs/$RUN_ID/cancel poll status_url until payload.status is canceled.
One more access rule: Agent Completions need agent_completions:write, which keys made before the feature shipped do not have. Those keys fail with 403 insufficient_scope, and the fix is a new key, since scopes cannot be added to an existing one. Service-account keys cannot create completions at all.
How to decide
Choose the three calls when the steps are fixed and you want control: you know the order, you can store ids, and each step has a clear price. Choose an Agent Completion when the steps depend on what the agent finds, and you want it to decide. Either way, put a ceiling on spend: the completion requires generation_spend_cap_usd, and your own code should check the balance before each paid step.
A completion call returns 202, and you poll /v1/agent-runs/$RUN_ID for the outcome. It needs the agent_completions:write scope on the key. You do not get a per-step price list in advance, so the cap is what protects you.
The three-call route gives you a log line for each paid step. The completion gives you one run to read. Pick the one that matches how you want to debug.
- Fixed order: three calls.
- Open-ended work: Agent Completion.
- Always set a spend cap on completions.
Sources
Related posts
More in Agents
- Agent Completions input: JSON data the agent reads from a file
Send caller data in input, not in instruction. Sume writes it to /workspace/inputs/sume-action-input.json and tells the agent to read it as data. curl sample.
- Agent Completions request limits: 100,000 characters, 50 messages
The Sume Agent Completions schema caps instruction at 100,000 characters, messages at 50, and images at 30. Know where each limit is enforced before you send.
- Preflight a paid Sume MCP call: balance_get, then dry_run
Before a paid Sume tool call, an agent should read the balance and run the tool with dry_run=true. Then submit with a fresh idempotency_key and an optional cap.
- API-triggered scheduled run: key per week, not $(uuidgen)
The docs' example sends Idempotency-Key: $(uuidgen), which changes on every retry. Derive the key from the week so a retry returns the same run.
Written by Sume