On a 402, try a cheaper rung: a Python ladder with fresh keys
A 402 on Sume means nothing was reserved, so a cheaper request can go straight through. Python ladder: Seedance 720p, 480p, then Wan 480p, one key per rung.

When a Sume video submit returns 402 insufficient_credits, you can retry the same prompt on a cheaper request immediately, because the 402 is returned before any provider work starts and no hold is taken. The one rule that trips people up is the idempotency key: a different payload under the same key returns 409 idempotency_conflict, so every rung of the ladder needs its own key.
The script below walks three rungs, Seedance 2.5 at 720p, Seedance 2.5 at 480p, then Wan 3.0 at 480p, and stops at the first one the balance can cover. The error codes come from Errors and rate limits and Generation admission.
Why a 402 is safe to step down from
Sume reserves the listed price times 1.25 when you submit. If the balance cannot cover that reservation, the API answers 402 with next_action: add_funds and retryable: false, and nothing is created. The retryable: false flag means retrying the identical request will fail again until funds are added; it does not stop you from sending a different, cheaper request.
That is the point of a ladder: you are not retrying, you are submitting a new request whose reservation is smaller. A top-up is dashboard-only, so the final rung of any script should stop and tell a human, not loop.
The ladder
Order the rungs by what you would rather keep. The prices below are the reservation for a 5 second 16:9 clip, list times 1.25 rounded up to the cent, from Sume's price tables on main.
| Rung | Request | Reserved |
|---|---|---|
| 0 | Seedance 2.5, 720p | $2.89 |
| 1 | Seedance 2.5, 480p | $1.35 |
| 2 | Wan 3.0, 480p | $0.32 |
import json, os, urllib.request, urllib.error
API = "https://api.sume.com/v1/videos"
LADDER = [
{"model": "seedance-2.5", "resolution": "720p", "duration": 5},
{"model": "seedance-2.5", "resolution": "480p", "duration": 5},
{"model": "wan-3.0", "resolution": "480p", "duration": 5},
]
def submit(body, key):
req = urllib.request.Request(API, json.dumps(body).encode(), {
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json", "Idempotency-Key": key})
with urllib.request.urlopen(req) as r:
return json.load(r)
def run(prompt, batch):
for i, rung in enumerate(LADDER):
try:
return submit({**rung, "prompt": prompt}, f"{batch}-rung{i}")
except urllib.error.HTTPError as e:
if e.code != 402:
raise
raise SystemExit("Out of funds: top up in the dashboard")
print(run("A paper boat on a rain puddle", "demo-001"))What each design choice does
Each rung gets its own Idempotency-Key, built from the batch name and the rung index. If your process crashes and restarts, it replays the same keys: a rung that was accepted returns the original job instead of creating a second paid one.
Only 402 steps down. A 429 queue_full or rate_limited is not a funding problem and a cheaper request will not help; wait and retry with the same key, as the docs say. A 400 or 401 is a bug in the request and should be raised, which is why the script re-raises anything that is not a 402.
- Never reuse a key across rungs: the payload differs, so the second call returns 409.
- Reuse the same key for the same rung after a crash: the replay is safe.
- Stop the ladder at the last rung instead of looping, since only the dashboard can add funds.
Choosing rungs that mean something
A ladder is only useful if each step is acceptable to the person who asked for the clip. Dropping 720p to 480p changes sharpness; dropping from Seedance to Wan changes the look and the model behavior. For drafts, that trade is usually fine. For a final, you may prefer the script to stop and notify instead of quietly delivering a different model.
Record which rung succeeded next to the job. The response carries the model you asked for, so storing the rung index with it makes cost reviews simple: you can see how often the wallet forced a downgrade and decide whether to raise the funded balance rather than keep lowering quality. The Usage page shows reserved, captured and refunded amounts, which confirms that skipped rungs cost nothing.
If you submit many clips, add a small guard: after the first downgrade in a batch, start the remaining clips at the rung that worked instead of trying the top rung again. The balance has not grown, so every extra top-rung attempt will just return another 402. Reading GET /v1/balance once per batch and comparing it with the rung prices is cheaper still, and it lets you plan the whole run before the first submit.
Sources
Related posts
More in Developers
- Sume failed job says [redacted_url]: what was removed
A [redacted_url] in a Sume job error is by design: URLs, provider ids, env names and secrets are masked, and the provider reason is capped at 300 characters.
- Is there a lipsync-1.0 endpoint on Sume? Old paths 404; use Fabric
Sume's old /v1/lipsync-1.0 paths return 404 and its model ids return model_not_found. Send the same still and audio to veed/fabric-1.0 or H3 Max lip-sync.
- Make a vertical Short clip with curl and jq: submit, poll, download
A 20-line shell script that asks Sume for an 8-second 9:16 clip at 1080p, polls the job, and saves short.mp4, matched to YouTube's Shorts page.
- Node script for a 9:16 TikTok video: check the model, then submit
A Node 20 fetch script that confirms a Sume model lists 9:16 and your duration, submits one 12-second 720p job, and saves an MP4 that fits TikTok's API limits.
Written by Sume