Four music variants in parallel: Python asyncio, one key per take
Submit four Music Router jobs at once from Python with one Idempotency-Key each, poll status concurrently, and pay 4 x $0.125 = $0.50. Code that runs.

To get four Music Router variants at once, run four submit-and-poll coroutines with asyncio, each with its own Idempotency-Key. The cost is 4 x $0.125 = $0.50, and a retry with the same key does not create a second job. The script below uses only the standard library.
The script
It posts to /v1/music-router/generate with sume/music-auto, then polls /v1/jobs/{id}/status until terminal is true, then reads /v1/jobs/{id}/result. Calls to urllib run in threads so the four jobs overlap. Set SUME_API_KEY first.
import asyncio, json, os, urllib.request
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"}
def call(method, path, body=None, key=None):
h = dict(H, **({"Idempotency-Key": key} if key else {}))
data = json.dumps(body).encode() if body else None
req = urllib.request.Request("https://api.sume.com" + path, data, h, method=method)
with urllib.request.urlopen(req) as r:
return json.load(r)
async def take(n, bpm):
prompt = f"Warm lo-fi, {bpm} BPM, C minor. A 30-second track. Instrumental, no vocals."
job = await asyncio.to_thread(call, "POST", "/v1/music-router/generate",
{"model": "sume/music-auto", "prompt": prompt}, f"jingle-take-{n}")
jid = job["request_id"]
while True:
st = await asyncio.to_thread(call, "GET", f"/v1/jobs/{jid}/status")
if st.get("terminal"):
break
await asyncio.sleep(st.get("next_poll_after_seconds") or 5)
return await asyncio.to_thread(call, "GET", f"/v1/jobs/{jid}/result")
async def main():
results = await asyncio.gather(*(take(i, 72 + 12 * i) for i in range(4)))
print(json.dumps(results, indent=2)[:2000])
asyncio.run(main())Why each take has its own key
Writes need an Idempotency-Key. If you reuse one key across four different prompts, the later requests can be treated as repeats of the first. A key per take (here jingle-take-0 to jingle-take-3) means a network retry returns the same job, and a deliberate new variant uses a new key.
The tempos step 12 BPM apart (72, 84, 96, 108), which the Music docs suggest when you want clearly contrasting takes.
Reading the result
Read the audio artifact from result.artifacts where type is audio. When present, result.lyrics carries the model-reported lyrics or section map. job.request.routed_model tells you which engine ran. The prompt asks for a 30-second track, but there is no duration parameter, so check the length of what comes back and trim with Timeline if needed.
Extending it
To run twelve takes, change range(4) to range(12), but respect your account's concurrency limits: if you hit a rate limit, run in batches of four. A twelve-take run costs 12 x $0.125 = $1.50.
The script prints the first 2,000 characters of the combined JSON. In production, pull the audio URL from the artifacts list and download it right away, since durable media URLs are what you want to keep.
Sources
Related posts
More in Developers
- Free plan accepts 6 paid jobs: submit 7 and read queue_full
Free allows 1 processing job and 5 queued, so 6 are accepted. A Python script submits 7 in parallel and prints which are accepted and which get 429 queue_full.
- Free, Pro, Startup, Scale: processing seats, queue slots, full hold
Sume's concurrency by plan, queue capacity max(3, 5 x concurrency), accepted job capacity, and the balance reserved if every slot holds a 10 s clip.
- Gate a Format run on TikTok limits: artifact size, length, pixels
Check a Sume Format run's artifacts[] for size, duration and dimensions against TikTok's non-Spark limits before upload. A Python gate under 30 lines.
- gemini-nano-banana-2.1 or google/nano-banana-2.1: which id Sume takes
Google's pricing page names the model gemini-nano-banana-2.1. Sume's image API takes google/nano-banana-2.1 and runs the old nano-banana-2 as 2.1.
Written by Sume