Six locales, six Agent Completions at $2 each: a $12 worst case
One Agent Completion per locale, each with its own spend cap and Idempotency-Key. Six runs at $2 bound generation to $12, and a retry cannot start a seventh.

To produce one asset in six locales with the Sume Agent, start six Agent Completions rather than one that loops. Each gets its own generation_spend_cap_usd and its own Idempotency-Key, so a failure or retry in one locale cannot spend the other five locales' money. Six runs at a $2 cap each bound generation to 6 x $2 = $12.
Why one completion per locale
generation_spend_cap_usd has no default on Agent Completions, and the request fails with 400 invalid_request without it. It is a limit for that one run. A single run asked to do six locales shares one cap among them, so a bad first locale can use most of it.
Separate runs also fail separately. Each returns its own receipt and its own terminal webhook, and the receipt shows usage for that run only. You can then re-run just the locale that failed.
There is one more reason. An Agent Completion starts in a new thread and cannot continue a prior one: thread_id continuation and assistant turns in messages are not available. Six independent instructions fit that model better than one long conversation that you cannot resume if it fails halfway.
The numbers
The cap is a ceiling, not the expected cost. The table shows the ceiling for the batch and what a retry does when keys are set per locale, from the Sume docs read 2026-10-09.
| Item | Value |
|---|---|
| Runs | 6 (en, ko, ja, es, de, fr) |
| Cap per run | $2.00 |
| Batch generation ceiling | 6 x $2.00 = $12.00 |
| Idempotency-Key | order id plus locale, 1-255 characters |
| Same key, same body | Returns the original receipt, idempotency_hit true |
| Same key, different body | 409 idempotency_conflict |
| Agent LLM turn | Billed to the Agent wallet, not in the cap |
A starter script
The script starts the six runs. The key is the order id joined to the locale, so running the script twice after a crash replays receipts instead of creating runs 7 to 12. It needs a key created after Agent Completions shipped, with agent_completions:write.
import os, uuid, json, urllib.request
KEY = os.environ["SUME_API_KEY"]
LOCALES = ["en", "ko", "ja", "es", "de", "fr"]
CAP = 2 # USD per run; 6 x 2 = 12 worst case
def start(locale, order_id):
body = {
"instruction": f"Write a 20 word product teaser in {locale}.",
"generation_spend_cap_usd": CAP,
"communication": {"webhook_url": "https://example.com/sume-hook"},
}
req = urllib.request.Request(
"https://api.sume.com/v1/agent/completions",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
"Idempotency-Key": f"{order_id}-{locale}",
},
)
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["data"]["id"]
ids = [start(loc, "order-1042") for loc in LOCALES]
print(ids, "worst case:", CAP * len(LOCALES))Collecting the results
Poll each status_url, or accept the webhook. The webhook sends one signed POST for each run when it completes or fails, with the same receipt as the poll endpoint. Dedupe on the envelope request_id, which equals run_id. Branch on outcome, not on status, because a degraded outcome means a run completed and billed but produced no structured output.
A canceled run sends no webhook, so if you cancel a locale, poll that one run until it reads canceled. The ceiling in the table holds even if every locale hits its cap, which is what makes it safe to leave running unattended.
If the batch must fit a smaller budget, lower the cap rather than the locale count. Cap times runs is the only arithmetic you need, so a $1.50 cap gives 6 x $1.50 = $9.00. Sizing the cap to what one locale needs is a decision for you, and the API pricing page lists the metered rates a run can use.
Sources
Related posts
More in Agents
- Canceled and skipped Sume agent runs send no webhook: what to poll
A Sume run webhook fires only when a run completes or fails. Canceled and skipped runs send nothing, so your scheduler must read status_url for them.
- A 20,000-character narration fits a $1.00 scheduled cap at $0.95
Sume TTS at $0.0475 per 1,000 characters prices 20,000 characters at $0.95, 5 cents under the $1.00 default cap of a schedule. What 21,000 does.
- Unattended agent: stop or retry on Sume 402, 409, 429, 503?
An unattended Sume agent should stop on 402, fix on 400, retry 429 and 503 with the same idempotency key, and never reuse a key after a 409.
- Wan 3.0 clip lengths that fit a $1.00 scheduled run cap
A schedule without its own cap gets $1.00 per run. At Wan 3.0 rates that buys 16 s at 480p, 8 s at 720p or 4 s at 1080p; here is the table.
Written by Sume