LangGraph 1.2 node timeout: the Sume video job keeps billing

LangGraph 1.2 adds run_timeout and idle_timeout per node. A timeout stops your node, not the Sume job it started: store the job id and re-poll.

5 min readSume
All posts

A LangGraph node timeout stops your node, not the Sume job it started. Set run_timeout or idle_timeout on a node that calls Sume, keep the job id outside the timed-out attempt, and re-poll Sume's status endpoint instead of submitting again.

LangGraph v1.2.0 (May 12, 2026, per the LangChain Python changelog read 2026-10-02) lets you pass timeout= to add_node(). A TimeoutPolicy supports a wall-clock limit (run_timeout) and an idle limit (idle_timeout), and expiry raises NodeTimeoutError. The same notes say timeouts are Python-only and apply to async nodes only.

What does the timeout actually cancel?

Only the node attempt inside your process. A Sume job is a durable record on Sume's side: Jobs and results says a client-side timeout does not cancel the job, which keeps running and still bills, and that you have only stopped watching. Cancellation exists (POST /v1/jobs/{id}/cancel) but succeeds only before generation starts; after that the API returns 409 job_generation_already_started.

That makes a naive pattern expensive: a node submits a video, the 120-second run_timeout fires, the retry policy runs the node again, and the second attempt submits a second paid job.

Which timeout fits which part of a Sume call?

Split the work. Submit is fast; waiting is slow. Put the submit in one short node with a tight run_timeout, and the wait in a second node, so a timeout on the wait never repeats the submit.

Timeout choice for a Sume job (read 2026-10-02; LangGraph terms from the changelog)
NodeSuggested limitWhy
submit (POST with Idempotency-Key)run_timeout of tens of secondsThe response is a 2xx with a job id, not the finished video
wait (poll status)run_timeout above your own deadline, or idle_timeoutPolling returns a snapshot each time; sleep is your code, not Sume
fetch resultshort run_timeoutResult is read only once result_ready is true

How do I keep the job id across a timed-out attempt?

Write the job id to graph state in the submit node and make the idempotency key a function of the business item (order-8823-hero-v1), not the attempt number. If the submit node itself times out after the request was sent, retrying with the same key returns the original job instead of billing a second one; 409 idempotency_conflict appears only when the key is reused for a different payload (Generation admission).

The helper below is a plain async function you can call from a node. It submits with a stable key and polls with next_poll_after_seconds, giving up with the job id so the graph can resume later.

import asyncio, os, time
import httpx

BASE = "https://api.sume.com/v1"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

async def render(prompt: str, key: str, deadline_s: float = 600) -> dict:
    async with httpx.AsyncClient(headers=AUTH, timeout=40) as http:
        r = await http.post(f"{BASE}/image-1.0/generate",
            headers={"Idempotency-Key": key},
            json={"prompt": prompt, "mode": "async"})
        r.raise_for_status()
        job_id = r.json()["request_id"]
        end = time.monotonic() + deadline_s
        while time.monotonic() < end:
            s = (await http.get(f"{BASE}/jobs/{job_id}/status")).json()
            if s.get("terminal"):
                return {"job_id": job_id, "status": s.get("sume_status")}
            await asyncio.sleep(s.get("next_poll_after_seconds") or 5)
        return {"job_id": job_id, "status": "still_running"}

async def main():
    print(await render("matte black bottle on marble", "order-8823-hero-v1"))

asyncio.run(main())

How do I test the timeout path without paying twice?

Run the graph with a run_timeout shorter than a normal render, and check three things in your logs: the state holds one job_id after the timeout, the second attempt sent the same Idempotency-Key, and the Sume dashboard shows one job for the item. If it shows two, the key is being built from something that changes per attempt.

Also test the opposite case, where the job finishes while your node is timing out. The next poll should see terminal: true and result_ready: true and fetch the result once. Reading the result earlier returns 409 job_not_completed, which is a signal to wait, not a failure.

Last, test a failed job. Sume gives it a public error on the job record, and your graph should route it to a repair or human step rather than the timeout handler. A timeout is about your clock; failed is about Sume's outcome, and mixing them hides real provider errors behind a retry loop that bills each time.

What if the job is still running when my deadline hits?

Return still_running with the job id, route the graph to a later node or an interrupt, and poll again from state. Do not call it failed: failure is a Sume status (failed, with a public error), not your deadline. If you really want to stop spend, cancel before it starts, knowing the cancel may be refused.

Sume has no event stream; GET /v1/jobs/{id}/events is a pull snapshot, and job webhooks only fire on terminal events. For TypeScript graphs the SDK ships waitForJob, which is this same loop (Runs and waiting), though note LangGraph's own node timeouts are Python-only per the changelog.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume