Should my backend call Sume over hosted MCP or the REST API?
REST from a backend, hosted MCP from an agent client. Where they differ: auth, wait limits, REST-only Image 1.0 and Video 1.0, and write budgets.

Call the REST API at https://api.sume.com/v1 from your own backend, scripts and queue workers, and use hosted MCP at https://mcp.sume.com/mcp when the caller is an agent client such as Claude Code, Cursor or Codex that needs to discover tools by itself. The docs call hosted MCP the preferred MCP path, but they place backend integrations on the Developer API. Two products make the choice easy: Image 1.0 and Video 1.0 are REST-only.
The comparison
Both surfaces create the same kind of durable jobs, bill the same wallet, and use the same job ids. What changes is who is calling and how the caller learns the outcome.
| Question | Hosted MCP | REST API |
|---|---|---|
| Endpoint | https://mcp.sume.com/mcp | https://api.sume.com/v1 |
| Typical caller | An agent client that lists and calls tools | Your backend, scripts, queue workers |
| Auth | OAuth with mcp:read by default and optional mcp:write, or an API key | API key as Authorization: Bearer or x-api-key, one at a time |
| Paid and write calls | Need an idempotency_key argument | Send an Idempotency-Key header |
| Image 1.0, Video 1.0 | Not available (images_create and videos_create stay REST-only) | /v1/image-1.0/generate, /v1/video-1.0/generate |
| Waiting for a result | jobs_wait in slices: default 50 s, cap 55 s, then call it again | Poll status_url, or sync for at most 30 s, or a webhook |
| Discovery | tools_list, tools_schema, catalog_list | GET /v1/catalog and the OpenAPI document |
Auth is the sharpest difference
Hosted OAuth is read-only unless the person at the consent page turns Write on, and there is no mcp:paid scope: paid submits are gated by wallet and admission, not by a scope. A session with only mcp:read can list jobs and read assets but gets insufficient_scope on generate_image. An API key on the MCP endpoint sees the full tool set. An OAuth token is not an API key and an API key is not an OAuth token, so you cannot swap one for the other.
A backend has no person to click a consent page, so it uses an API key and the REST API. An interactive agent in a developer's editor has a person, so OAuth is the better fit.
Waiting and delivery
The REST API offers four communication modes: async (the default), sync and its alias subscribe (one bounded wait, at most 30 seconds, then poll), and webhook. For a backend that handles many jobs, webhook plus a poll fallback is the production shape.
Over MCP, a single jobs_wait call holds for at most 55 seconds, because an HTTP request that stays open longer is killed at the edge. For a ten-minute render, call it again with the same ids; do not submit the paid create again. jobs_wait also takes up to 20 ids and a wait_for of all or any, so one call can cover a whole fan-out.
Rate limits differ in a useful way
An MCP tool call spends the write budget once, for the run it creates, and not for the JSON-RPC request that carried it. A jobs_status poll over MCP spends none. If your REST backend is close to its write budget, that is a point in favor of letting an agent drive the long polling over MCP; but a backend you control can also poll cheaply, because REST reads sit in their own bucket at forty times the write number.
Idempotency on both sides
The rule is the same on both surfaces and only the spelling changes. On REST you send the Idempotency-Key header, and on hosted MCP you pass idempotency_key in the tool arguments. Both exist so a retry after a timeout returns the original job instead of billing a second one. Over MCP the key is described as transport and dedup, not human approval, so a clean agent loop generates one stable key per intended render and reuses it on retry.
If your agent runs long loops, give it the spend gates too. dry_run previews cost without submitting, generation_admission_preview shows the queue and balance, and max_spend_usd caps a call when you set it. On REST the same protection comes from GET /v1/balance, the generation_limits object on submit responses, and the per-run generation_spend_cap_usd on Formats.
A rule of thumb
If the thing that decides what to generate is your code, use REST. If the thing that decides is a model choosing among tools at run time, use hosted MCP and let it call tools_list and tools_schema for the live contract. If you need both, they can share one workspace and one wallet, and the job ids you get from one are readable from the other under the same key's visibility rules.
Sources
Related posts
More in Developers
- Speech to text API in Go: transcribe audio with net/http
Transcribe audio in Go using only the standard library: submit to Sume STT, poll the job and print the text. A 30-line program at one cent per audio minute.
- Speech to text API in Node.js: transcribe audio with fetch
Transcribe audio in Node.js with built-in fetch: submit to Sume STT, poll the job and print sentence segments with timestamps. 30 lines, no dependencies.
- Speech to text API in Ruby: transcribe audio with Net::HTTP
Transcribe audio in Ruby with only the standard library: submit to Sume STT, poll the job, print text and word times. A 30-line script at one cent a minute.
- Speech to text API in Swift: transcribe audio with URLSession
Transcribe audio in Swift with async URLSession: submit to Sume STT, poll the job, print text. No packages, 28 lines, one cent per audio minute of audio.
Written by Sume