AI SDK stream cancel on disconnect: the Sume job keeps running
AI SDK 7.0.127 fixes stream cancellation when consumers disconnect. A cancelled stream does not cancel a Sume job: save the job id and read status later.

When an AI SDK stream is cancelled because the browser tab closed, a Sume job your tool already submitted keeps running and keeps billing. Save the job id the moment submit returns, and let the next request read GET /v1/jobs/{id}/status instead of submitting again.
The vercel/ai releases page (read 2026-10-02) lists "fixed stream cancellation when consumers disconnect" under ai@7.0.127. That is a good fix for your server: it can stop work when nobody is listening. It cannot reach into Sume, which has its own job lifecycle.
What does a cancelled stream stop?
Your route handler and anything it awaits. A tool that called Sume with mode: "async" got a job id in the first response, because Sume accepts a submit the moment it has a durable job id. Everything after that is Sume's state. Jobs and results says a client-side timeout does not cancel the job, and the same holds for a client that left.
So a user closes the chat mid-render, the stream aborts, and the job completes into the void, charged. If the user returns and asks again, a model that submits a fresh request pays twice.
Where should the job id live?
In your database, written inside the tool before it returns, keyed by the conversation and the item. The tool's output also goes into the message history, so the model sees the id on the next turn. Use a deterministic Idempotency-Key so a retried tool call after a reconnect returns the original job.
curl -X POST https://api.sume.com/v1/image-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: chat-77-msg-3-hero" \
-d '{"prompt":"Product hero shot on marble","mode":"async"}'
curl https://api.sume.com/v1/jobs/job_123/status \
-H "Authorization: Bearer $SUME_API_KEY"Should the disconnect handler cancel the job?
Only if you accept it may be refused. POST /v1/jobs/{id}/cancel succeeds before generation work starts; after that it returns 409 job_generation_already_started with details.cancelable: false, and the job runs to completion. A queued job in a busy workspace is cancelable, a processing one usually is not.
| Policy | Result | Trade-off |
|---|---|---|
| Do nothing | Job completes and bills | Result waiting when the user returns |
| Cancel on disconnect | Works only before generation starts | May return 409; job still bills if it started |
| Switch to webhook mode | Callback stores result | Needs a public HTTPS endpoint and signature check |
How do I test a disconnect?
Start a render from the UI, close the tab right after the tool returns, and look at three places. Your database should hold one job id for the chat. Sume should show the job moving through queued, processing and a terminal state. When you reopen the chat, the UI should find the job id and show the result or the running status, with no new submit.
Repeat with a flaky network: kill the connection after the request leaves but before the response arrives. The retried tool call must send the same Idempotency-Key. If your dashboard shows two jobs for one chat message, the key is built from something that changes per attempt.
Finally, decide who sees the result when the user never returns. A job that completes with nobody watching still bills. A webhook that stores the artifact URL against the chat means the work is waiting when they come back, and an unwatched job is at least not wasted.
What about keeping the stream alive instead?
For short image jobs you may be able to hold the stream. For video, hold nothing: Sume's sync mode waits at most 30 seconds, and its own docs call 30 seconds a wait budget, not a job duration. Submit async, tell the user it is running, and resume from state. See AI SDK SSE heartbeat vs Sume jobs_wait slices for the in-stream variant. Before you ship, read the live contract for every route you call at https://api.sume.com/reference/json, which Sume's docs name as the schema source of truth, and re-read the linked docs pages: limits, scopes and error codes change faster than blog posts do. Treat any number in this post as a snapshot dated 2026-10-02, and prefer the effective fields your own responses return, such as generation_limits, over a static table.
Sources
Related posts
More in Developers
- AI SDK 7 tool search with deferred tools and Sume tools
AI SDK 7.0.127 lets a search() callback rank eligible deferred tools. Load a few Sume tools up front and fetch the rest by tools_schema on demand.
- Did the edit stay in its region? A pixel-diff check for GPT Image 2.5
After a GPT Image 2.5 edit, measure how much changed outside the area you meant to change. A Pillow script that diffs the result against the original.
- Verify a Sume avatar video webhook in Python (HMAC SHA-256)
A Python verifier for avatar video webhooks: timestamp tolerance, rotation-safe comparison, and a hard refusal when the signing secret is empty.
- low_confidence_long_video: why video_frames warns past 90 seconds
Sume's video_frames returns the low_confidence_long_video warning when the source runs over 90 s. The job still succeeds; the hard cap is 300 s. What to do.
Written by Sume