Recast request timed out: retry with the same Idempotency-Key
A timeout on POST /v1/video-router/generate does not tell you whether the job exists. Derive a stable Idempotency-Key so a retry returns the original job.

If a Recast submit times out, retry it with the exact same Idempotency-Key and Sume returns the original job instead of creating and charging a second one. The docs say "send Idempotency-Key to make retries safe; a replay returns the original job" (Video generation docs, read 2026-10-03). Without a key, a network timeout leaves you unable to tell whether the server received the request, and a blind retry can double your spend on a clip that costs several dollars.
The practical rule is to compute the key from the content of the request, not from a clock or a random number, so the same logical job always carries the same key.
Why a timeout is ambiguous
A client timeout means your socket gave up, not that the server did nothing. The request may have been dropped on the way in, accepted and answered after you stopped listening, or half-processed. Sume accepts a submit the moment it has a durable job id, so in the middle case a job exists that you have no id for. A fresh submit with a fresh key would then create a second job and reserve the cost twice, which at $5.63 for a 15-second 768p clip is small once and large across a batch.
With the same key, both outcomes converge on one job. If the first request did arrive, the retry returns that job. If it did not, the retry creates it.
Deriving the key
Hash the things that define the job: the model, the source URL, the sorted photo URLs and the resolution. If you change any of them you intend a different job, and the key changes. Add a version suffix you control for the case where you deliberately want to rerun the identical request, since Recast takes no seed and a rerun is how you get another take.
import asyncio
import hashlib
import json
import os
import httpx
def job_key(source: str, photos: list[str], resolution: str, rerun: int = 1) -> str:
raw = json.dumps([source, sorted(photos), resolution, rerun])
return "recast-" + hashlib.sha256(raw.encode()).hexdigest()[:24]
async def submit(source: str, photos: list[str], resolution: str = "768p") -> dict:
key = os.environ.get("SUME_API_KEY")
if not key:
raise SystemExit("set SUME_API_KEY")
body = {"model": "h3-max-recast", "video_url": source,
"reference_image_urls": photos, "resolution": resolution}
headers = {"Authorization": f"Bearer {key}",
"Idempotency-Key": job_key(source, photos, resolution)}
async with httpx.AsyncClient(timeout=20) as c:
for attempt in range(3):
try:
r = await c.post("https://api.sume.com/v1/video-router/generate",
json=body, headers=headers)
r.raise_for_status()
return r.json()
except httpx.TransportError:
await asyncio.sleep(2 ** attempt)
raise RuntimeError("submit failed after 3 attempts")
print(asyncio.run(submit("https://example.com/a.mp4", ["https://example.com/p.jpg"])))A worked failure
Suppose you submit a 15-second 1080p Recast from a worker that has a 10-second HTTP timeout. The submit leaves your machine at second 0. Sume accepts it, creates the job and reserves about $8.44. At second 10 your client gives up. If your code now generates a fresh random key and retries, the second submit creates a second job and reserves another $8.44. Both run to completion; you download one and leave the other unused. The money is spent twice.
With a content-derived key the second submit returns the first job. You get one job, one hold and one clip. This is the entire reason the header exists, and it costs a single line in a client.
What not to retry
Retry transport failures and server errors. Do not retry a 400: the request is wrong and the same key will produce the same answer. Do not retry on a 402 or 429 without waiting and checking your balance or queue, which the generation admission docs describe. And do not reuse a key for a request whose body you changed; use a new key, because the point of the key is that it names one logical job.
| Result | Retry? | How |
|---|---|---|
| Socket timeout or reset | Yes | Same key, short backoff |
5xx | Yes | Same key, backoff |
400 (bad field, shot over 15 s) | No | Fix the request, then use a new key |
402 or 429 | After checking | Same key once balance or queue allows |
Job failed after acceptance | Judgement | Failed jobs are refunded; resubmit with rerun bumped |
After the retry
Once you hold a job id, stop submitting and start following it. Poll GET /v1/jobs/{id}/status or use a webhook, as the Jobs and results guide describes. Store the key beside the job id in your own records so that a support request or a duplicate-spend check can match the two. If a clip is rejected by review and you want another take, bump the rerun counter, which is a new key and a new, deliberate charge.
Sources
Related posts
More in Developers
- Recast sync wait caps at 30 seconds: use a webhook for long sources
Sync and subscribe modes wait at most 30 seconds, so submit h3-max-recast jobs as async or webhook and poll /v1/jobs/{id}/status as a fallback.
- Recraft V4 on Sume returns WebP only: convert to PNG or JPEG in Python
Recraft V4 on Sume outputs WebP and takes no references. A Pillow converter for PNG or JPEG, with transparency flattened onto white for JPEG delivery.
- Reel judders after mixing 24 and 30 fps clips: set Timeline output fps
Timeline repeats or drops frames when output fps differs from a source. Learn output_fps_resamples_sources, how the default is chosen, and when to pin 30.
- reference_video_urls or video_url? Reference footage vs edit source
On Sume, reference_video_urls guide a new clip; video_url is a source you edit or swap. They cannot be combined on Gemini Omni Flash. Which model takes which.
Written by Sume