MCP TS SDK 2.3 enforces one server per request: where state lives

TypeScript SDK 2.3.0 enforces one server instance per request. For a media tool that means job state belongs in job ids, as Sume's jobs_wait does.

4 min readSume
All posts

The MCP TypeScript SDK v2.3.0 enforces one server instance per request, so nothing you keep in a server object survives to the next call. A video tool therefore has to put its state in a job id that the caller sends back, which is how Sume's hosted MCP tools already work.

This post covers what changed in the release and what a media-generation server should do about it.

What the 2.3.0 release lists

The 2026-07-28 MCP specification removes protocol-level sessions and the Mcp-Session-Id header from Streamable HTTP. Each request stands alone.

TypeScript SDK v2.3.0, released Oct 2, 2026 (read 2026-10-03)
ChangeWhat it means for a media tool
One server per request enforcedDo not hold a job map in module or instance memory.
Same-origin redirects by defaultPoint clients at the final MCP URL, not a redirecting alias.
50 MB SSE tool result now under a second (was about 13 s)Big results are cheaper, but still return asset URLs.
maxToolInputElementsCap array arguments, such as a list of scene prompts.
tasks/get and tasks/cancelPolling and cancel calls for long-running work.

Put the state in the id

A generation takes longer than a request. The usual way to bridge that gap without sessions is to return a durable identifier from the submit call and make every later call take that identifier as an argument.

Sume jobs work this way. Every generation endpoint creates a durable job and returns its id in the first response, in every communication mode. The docs tell clients to store the job id so an integration can recover work after a process restart.

On hosted MCP the follow-up tools take the id as an argument: jobs_status, jobs_result, jobs_events and jobs_wait. Nothing about the earlier call needs to be remembered by the server instance that handles the next one.

Waiting without a session

jobs_wait accepts a single job_id or a list of 1 to 20 job_ids with wait_for set to all or any. Each call holds at most 55 seconds, and the default is 50. When the slice ends with wait_slice_expired, call jobs_wait again with the same ids.

That retry is safe precisely because the job lives outside the connection. The docs are explicit that you must not resubmit the paid create, because the job keeps running and keeps billing.

  • Submit once with a stable idempotency_key; paid and write tools require it.
  • Keep the returned job ids in your own store, not in server memory.
  • Wait in slices and read jobs_result once the job is terminal.
  • Pass include_results: true to jobs_wait to get finished results in the same answer.

Checklist for your own server

If you are building a wrapper that exposes Sume generation as your own MCP tools, treat the SDK change as a prompt to audit three things.

  • No per-session maps: a submit tool should return the Sume job id and nothing else it needs later.
  • Idempotency belongs in the tool arguments, so a retried tool call after a dropped connection does not double-bill.
  • Return media.sume.com artifact URLs from results rather than inline media, as the job result shape already provides.

Related reading

The tools and gates page lists the full hosted tool inventory, and Jobs and results describes the status values and wait slices in detail.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume