Backend job that needs an agent: Agent Completions or hosted MCP?

Agent Completions runs Sume's agent and returns a 202 receipt with a required spend cap. Hosted MCP is for a client whose own model calls Sume tools.

4 min readSume
All posts

Pick by who supplies the model. If your backend wants Sume's own agent to do an ad-hoc task, call Agent Completions: POST /v1/agent/completions returns a 202 receipt, and generation_spend_cap_usd is required. If you want your own model (Claude, Mistral, anything else) to call Sume tools, connect it to hosted MCP at https://mcp.sume.com/mcp. The first runs Sume's agent; the second lets another agent use Sume.

Side by side

Both are documented on Sume's docs (read 2026-10-08).

Agent Completions versus hosted MCP, from Sume docs read 2026-10-08
QuestionAgent CompletionsHosted MCP
Whose model runsSume's agent; model is only sume-agentYours, in the MCP client
EntryPOST /v1/agent/completionshttps://mcp.sume.com/mcp
AuthAPI key with agent_completions:writeOAuth or API key
Response202 and an agent.run receipt you pollTool results, then jobs_wait
Spend controlgeneration_spend_cap_usd, required, no defaultdry_run, max_spend_usd (optional)
Sync chat wireNo; streaming not availableNot applicable

Details that bite

Agent Completions has constraints that a drop-in chat integration would not expect:

  • It is not a synchronous chat completion. The response is a run receipt with status_url and cancel_url, not choices[].
  • Send exactly one of instruction or messages. Assistant turns are rejected, not ignored, and each completion runs in a new thread.
  • Keys created before the feature shipped lack the scopes and get 403 insufficient_scope. Create a new key and rotate to it.
  • Service-account keys cannot create completions.
  • A repeated Idempotency-Key returns the original receipt with idempotency_hit: true; the same key with a different payload returns 409 idempotency_conflict.

A minimal call

The cap is the field to get right. Here the run may spend at most $2 on generation.

curl -sS -X POST https://api.sume.com/v1/agent/completions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: weekly-teaser-2026-10-08" \
  -d '{
    "instruction": "Make one 9:16 teaser clip for the new product.",
    "generation_spend_cap_usd": 2
  }'

When to choose which

Choose Agent Completions when the task changes on each call and you want Sume's sandbox and tools without running a model yourself. If the task is the same recipe with new inputs, Sume's docs point to Formats, and for a clock-driven job, to Scheduled. Choose hosted MCP when a person works inside Cursor, Claude Code or VS Code, or when you already run an agent loop and only need Sume's tools. Do not mix them up with the in-app Studio Agent, which is not an MCP connector.

Polling the receipt

After the 202, poll the status_url from the receipt until the run reaches a terminal state, and call the cancel_url if you need to stop it. Store the Idempotency-Key with your own record of the task, so a retry after a network error in your backend returns the original receipt, not a second run. Keep the spend cap in your own config, because it has no default and a missing value is rejected.

Choosing keys and scopes

Create a dedicated API key for the backend with only the agent_completions scopes it needs, and keep it in your secret store, not in a repo. If a run needs more than the cap you set, raise the cap on the next request on purpose; do not make it large by default. The cap is the spend control, so pick it per task.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume