Translate a pack into 8 languages in parallel: queue limits by plan

Eight Ideogram 4.5 edits fit the Pro queue (24 accepted jobs) but not Free (6), so 2 get 429 queue_full. Python thread pool with retry; cost is $0.60 at medium.

4 min readSume
All posts

Eight language versions of one pack are eight separate edit requests. On a Pro plan all eight are accepted at once, because the queue holds 24 jobs and 4 process at a time. On a Free plan only 6 are accepted, so a burst of eight returns 429 queue_full for the last two, and the fix is a small thread pool with a retry.

The plan limits

Sume documents these as processing concurrency, queue capacity and accepted jobs per plan (Generation admission). queue_full is separate from rate_limited, and 402 insufficient_credits is checked before any provider work starts.

Generation admission by plan (read 2026-10-05)
PlanProcessing at onceQueueAccepted jobs8 edits at once
Free1562 refused
Pro42024all accepted
Startup84048all accepted
Scale20100120all accepted

Cost

Ideogram 4.5 at the default medium quality is $0.075 billed per image on Sume, so eight languages are $0.60. At high they are $2.20. Failed or cancelled jobs are not billed.

Python: capped workers and a retry on 429

Cap the workers at your plan's processing number so you never fill the queue yourself, and wait before retrying on 429.

import os, time, requests
from concurrent.futures import ThreadPoolExecutor

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
SRC = "https://example.com/pack.png"
LANGS = ["German", "French", "Spanish", "Italian", "Dutch", "Polish", "Swedish", "Japanese"]

def one(lang):
    body = {
        "model": "ideogram/ideogram-v4.5",
        "prompt": f"Translate all text on the pack into {lang}. Keep layout, fonts, logo and colours.",
        "input_references": [{"image_url": {"url": SRC}}],
    }
    for wait in (2, 5, 15):
        r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=120)
        if r.status_code != 429:
            break
        time.sleep(wait)
    return lang, r.status_code, r.json()

with ThreadPoolExecutor(max_workers=4) as ex:
    for lang, code, b in ex.map(one, LANGS):
        print(lang, code, (b.get("data") or [{}])[0].get("url"))

Two things to handle

A request that takes longer than 30 seconds returns a 202 job envelope instead of the image, so the data list will be missing for that language. Poll the job as described in Jobs and results rather than counting it as a failure. Send an Idempotency-Key per language if your worker can retry after a timeout, so a retry does not buy a second image.

When not to parallelise

On Free, run one at a time. The queue is short and a retry loop only adds noise. And review each output before you publish: a language version that comes back with a changed ingredient line is cheaper to catch before it is in the catalogue.

Sizing the pool

Start with the plan's processing number as the worker count. Pro is 4, Startup is 8 and Scale is 20 in the table above. A larger pool only moves jobs from your side to Sume's queue, and on Free the queue holds just 5 behind the 1 being processed.

If you run several batches at once, they share the same account limits. Add the batches together before you decide, and keep the retry delays short on the first try so a refused job is not waiting in your code for long.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume