Music Router metadata: tag 12 candidate tracks for $1.50 in total
Music Router metadata is stored on the job, not sent to the provider. Tag 12 candidates at $0.125 each, $1.50, and read routed_model to see which engine ran.

The metadata field on a Music Router request is caller data that Sume stores on the job. Sume does not send it to the provider, so it cannot steer the music, but it lets you label a batch of candidates and find them again. Twelve candidate tracks cost 12 x $0.125 = $1.50 at the fixed per-generation price.
What is stored and what is sent
Two facts from the docs matter here. First, metadata stays in Sume. Second, job.request.routed_model names the engine that actually ran, for example lyria-3.5, while job.model echoes the id you asked for (sume/music-auto stays sume/music-auto).
| Value | Goes to the provider | Where you read it |
|---|---|---|
| prompt | yes | job.request |
| metadata | no | stored on the job |
| model (requested id) | used for routing | job.model |
| engine that ran | not applicable | job.request.routed_model |
| image_url | yes, as visual conditioning | job.request |
A labelled candidate batch
A common use is a shortlist. You write one brief, send it 12 times with different labels such as the take number or the mood word you varied, and compare. Every request should carry its own Idempotency-Key, so that a retried call does not create a second paid track. The docs example uses a key per request.
Price the batch first. Each accepted generation is $0.125, and a prompt change does not change the price. 12 x 0.125 = 1.50, so the batch is $1.50. A batch of 24 would be $3.00. The pricing note in the docs says the same fixed price applies to every Music Router model, and the catalog lists the provider list price per model for reference only.
import json
BRIEF = "Warm lo-fi hip hop, 84 BPM. Dusty Rhodes chords. Instrumental, no vocals."
PRICE = 0.125
bodies = []
for take in range(1, 13):
bodies.append({
"headers": {"Idempotency-Key": f"lofi-shortlist-take-{take:02d}"},
"json": {
"model": "sume/music-auto",
"prompt": BRIEF,
"metadata": {"batch": "lofi-shortlist", "take": take},
},
})
print(len(bodies), "requests")
print("cost: $%.2f" % (len(bodies) * PRICE))
print(json.dumps(bodies[0], indent=2))Choosing labels that you can use later
Keep the metadata small and flat. A batch name, a take number and the variable you changed are enough. Do not put secrets or personal data in it, because it is stored on the job. Because the provider never sees it, you also cannot use it to pass style hints. If you want a different mood, change the prompt.
A good label scheme makes the shortlist usable a week later. Name the batch by date and brief, number the takes, and add one field for the single thing you varied, such as tempo or lead. When you pick take 7, you know that tempo 84 won, and the next batch can start from that. Without labels you only have a list of job ids.
A note on cost control
Because a generation is billed when it is accepted, the cheap way to run a shortlist is small batches. Send three takes, listen, and only then send nine more. A shortlist of 12 at $1.50 is cheap, but 12 batches of 12 is $18.00. The metadata label keeps the batches apart so that you do not repeat a take that you already rejected.
Read it back
After a job finishes, GET /v1/jobs/{id}/result returns the audio artifact in result.artifacts[] where type is audio, and result.lyrics carries the model-reported lyrics or section map when they are present. Join those results to your own labels through the job id. If you pin lyria-3.5 or lyria-3-pro instead of sume/music-auto, compare routed_model on each job to confirm what ran.
Sources
Related posts
More in Media tools
- Music Router negative_prompt: send an empty string, not a list
Sume Music Router takes negative_prompt only omitted or as an empty string. Put exclusions in the prompt. Duration is refused too. Price stays $0.125.
- One 1920x1080 master to TikTok 9:16, 1:1 and 16:9 ads: $0.04
TikTok's page lists three ratios with minimum sizes. From one 1920x1080 clip, two video-filter crops make 9:16 and 1:1 at $0.02 each. The fractions are shown.
- One portrait, three languages: three 6-second lip-sync jobs for $1.80
A greeting in English, Spanish and Polish from one still: three 6-second H3 Max 768p jobs cost $1.80 plus TTS. Steps, the language check and limits.
- One TTS job per sentence, then a gapless concat: the Sume recipe
Generate narration one sentence at a time with tts_create, then join the takes with timeline audio. Fix one line without redoing the rest.
Written by Sume