GraphQL mutation for a Sume video submit: return the job id
Wrap a Sume video submit in a GraphQL mutation that returns the job id at once, then expose status as a query. No resolver should wait on a clip.

Put the Sume submit behind a GraphQL mutation that returns a job id and a status immediately, and put polling behind a separate query the client calls on its own schedule. A resolver that waits for a clip to finish ties a GraphQL request to a generation that can run for minutes, while Sume's bounded sync and subscribe waits stop at 30 seconds. The shape that survives is: mutation creates, query reads, webhook or poll finishes.
Why a resolver should not wait
GraphQL gateways, edge proxies and mobile clients have their own timeouts, and none of them know that a video job is still healthy. The jobs and results page is explicit that sync and subscribe are the same bounded wait of at most 30 seconds, and that a timed-out wait does not cancel the job. Treat that as a hint about the right API shape: a request you can finish quickly, and a handle you can read later.
The /v1/videos surface already works that way. POST /v1/videos answers 202 with an id, a polling_url, a status of pending, in_progress, completed, failed or cancelled, and the model. Your mutation can pass that triple through unchanged, so the client sees what Sume saw.
A mutation and a query
The sketch below is a resolver map with the HTTP client injected, so you can test it without a key. The Idempotency-Key is built from your own order id and version, which means a user who double-taps the button or a client that retries the mutation gets the same job back instead of a second paid one. Auth is one credential: the example sends x-api-key only, because sending both Authorization and x-api-key returns 401 unauthorized.
It runs as an ES module on Node 18 or newer (save it as a .mjs file) and prints the job id from the fake transport.
const BASE = "https://api.sume.com/v1";
export const resolvers = (fetchImpl, apiKey) => ({
Mutation: {
async createClip(_, { orderId, version, prompt }) {
const res = await fetchImpl(`${BASE}/videos`, {
method: "POST",
headers: {
"x-api-key": apiKey,
"content-type": "application/json",
"idempotency-key": `${orderId}:v${version}`,
},
body: JSON.stringify({ model: "sume/auto", prompt }),
});
const body = await res.json();
if (!res.ok) throw new Error(body.error?.code ?? `http_${res.status}`);
return { jobId: body.id, status: body.status };
},
},
Query: {
async clip(_, { jobId }) {
const res = await fetchImpl(`${BASE}/videos/${jobId}`, { headers: { "x-api-key": apiKey } });
const body = await res.json();
return { jobId, status: body.status };
},
},
});
const fake = async () => ({ ok: true, json: async () => ({ id: "job_demo", status: "pending" }) });
const r = resolvers(fake, "k");
console.log(await r.Mutation.createClip(null, { orderId: "A1", version: 2, prompt: "x" }));Map errors you can show
Do not expose Sume's message text to the client. Branch on the HTTP status and the code, then return a typed GraphQL error. A 4xx at create means nothing ran, so the mutation is safe to retry once the cause is fixed.
Credits deserve a dedicated error type: 402 insufficient_credits carries next_action: add_funds, and top-ups happen in the dashboard, so the client can show a message rather than an automatic retry.
| Sume response | GraphQL error code | Client behaviour |
|---|---|---|
| 400 invalid_request / unsupported_parameter | INVALID_INPUT | Fix the field named in details |
| 402 insufficient_credits | NEEDS_FUNDS | Show a top-up message; do not loop |
| 429 queue_full | BUSY | Wait for running jobs; keep the same key |
| 429 rate_limited | SLOW_DOWN | Back off with retry-after |
| 503 provider_capacity_exceeded | RETRY_LATER | Retry later with the same key |
Deliver the result without a subscription
You can offer a GraphQL subscription for UI convenience, but source it from a webhook or a server-side poll that honours next_poll_after_seconds; Sume itself has no SSE or WebSocket stream. Store the job id with the order, deduplicate completion events on it, and let the clip query read the stored state first. That keeps the read traffic of your own users off Sume's read budget.
Pagination and lists
If your schema exposes a list of clips, keep it in your own database and update it from webhook events or a sweeper, rather than resolving each item with a live call to Sume. A list of 50 rows that each triggers a status read is 50 reads per page view, and the read budget is shared with everything else on the key. Store status, result URL and timestamps on completion, and fall back to a single read only for rows that have been non-terminal longer than expected. That also makes the query fast, because it never leaves your datacentre.
Sources
Related posts
More in Developers
- H3 Max Recast job failed: what Sume refunds and what to retry
A failed Recast job releases its hold on Sume. Read the error category, fix input errors, retry queue errors with the same Idempotency-Key, never double-submit.
- H3 Max Recast seed: fal has one, Sume does not send it
fal's Recast API takes and returns a seed. Sume's Video Router accepts none, so each run is a new take. How to re-roll, what to vary, and what it costs.
- H3 Max Recast request shape: /v1/videos or /v1/video-router/generate?
Recast takes one video_url and 1 to 4 image_url entries as input_references on /v1/videos, or video_url plus reference_image_urls on the Video Router.
- H3 Max Recast webhook: submit with mode webhook, verify the HMAC
Recast jobs run for a while. Submit h3-max-recast with mode webhook, then verify Sume's sume-v1 HMAC signature before you download the swapped video.
Written by Sume