Design a render tool for a stateless MCP server: job ids as arguments
MCP 2026-07-28 removes protocol sessions. A render tool stays correct if its state lives in a job id the client passes back, as Sume jobs do.

Put the state in a job id, not in the connection. The MCP 2026-07-28 spec removes protocol-level sessions and the Mcp-Session-Id header from Streamable HTTP, and it removes the initialize handshake. A render tool that returns a durable job id and takes that id back as an argument works on any server instance. Sume's own job model has this shape, per its jobs docs.
What changed in the spec
Three items from the 2026-07-28 changelog, read on 2026-10-03, matter for a media tool.
With no session, two requests from the same client can reach different server processes. Anything you kept in memory between them is gone.
| Before | In 2026-07-28 |
|---|---|
| Session established by an initialize handshake | Handshake removed |
| Session tracked by the Mcp-Session-Id header | Header and protocol-level sessions removed |
| Version and capabilities negotiated once | Each request carries protocol version and client capabilities in _meta |
Three rules for a render tool
The rules follow directly from losing the session. Each one removes a place where a render could be lost or paid for twice.
- Return a job id from the submit call, and treat it as the only handle.
- Make every follow-up tool (status, wait, result, cancel) take that id as a required argument.
- Require a client-supplied idempotency key on the paid submit, so a retry after a dropped connection does not create a second render.
How Sume's job model fits
Sume jobs are durable and belong to a workspace and the member who created them, not to a connection. The docs say to store the job id from submit responses so an integration can recover work after a process restart, and not to resubmit a paid request because a local process timed out.
On the hosted MCP server, jobs_wait takes a job_id or up to 20 job_ids, holds for at most 55 seconds per call, and answers wait_slice_expired when the slice ends. The docs tell you to repeat the wait with the same ids. That loop needs nothing from a session: the ids are the state. Paid tools require an idempotency_key, described in the tools and gates page as a stable key for transport and dedup.
A tool contract you can copy
This is an example contract for a server you write yourself. It is not Sume's schema; read the real one with tools_schema. The point is that no tool depends on an earlier call having reached the same process.
{
"tools": [
{"name": "render_start",
"required": ["prompt", "idempotency_key"],
"returns": "job_id"},
{"name": "render_wait",
"required": ["job_ids"],
"note": "slice of 50 s or less; call again on expiry"},
{"name": "render_result",
"required": ["job_id"],
"returns": "artifact URLs, not inline video bytes"},
{"name": "render_cancel",
"required": ["job_id", "idempotency_key"]}
]
}Test it by killing the server
Run two instances behind a round-robin proxy and issue each call of a render to the other instance than the one before. If the flow completes, the tool is stateless in practice. If a call returns an unknown-job error, you kept something in process memory and it needs to move to a store keyed by job id.
Sources
Related posts
More in Developers
- Claude rejects forced tool_choice: steer generate_video by description
Claude Sonnet 5.5 and Opus 5.5 return a 400 for tool_choice any or tool, and thinking cannot be disabled. Steer Sume tool calls with descriptions and a dry run.
- Draft with GPT Image 2.5 Flare, finish with Sunburst: a two-pass edit
OpenAI pairs Flare with fast generation and Sunburst with editing precision. A Python two-pass on Sume's Image API that drafts, then refines the first result.
- Mcp-Name header rules: rate-limit paid render tools at the gateway
MCP 2026-07-28 requires Mcp-Method and Mcp-Name headers on Streamable HTTP POSTs. A gateway can rate-limit paid render tools by name without reading the body.
- GPT Image 2.5 on ElevenLabs: 14 ratios plus auto. Sume lists 17
ElevenLabs offers 14 fixed ratios plus auto for GPT Image 2.5. Sume's normalized list has 17 plus auto. Read what each model accepts before sending one.
Written by Sume