Music Router prompt limit: 5,000 characters, and what to put in them
Sume Music Router prompts run 1 to 5,000 characters, with no duration field or negative prompt; length goes in the text. A brief budget and a checker.

Short answer
A Music Router prompt can be 1 to 5,000 characters. That is far more than most briefs need, and the useful question is not how to fill it but what has to be in it, because the request has no other controls. Per the Music Router docs there is no duration field, no seed, no guidance setting and no accepted negative prompt.
Where each control lives
Everything that shapes the track goes in the one text field: length, structure, instruments and exclusions. The docs list three rules.
- Length: write it in the prompt, for example "a 2-minute track" or section markers such as
[0:00-0:30] Intro: .... Sendingdurationorduration_secondsis rejected. - Exclusions: put them in the positive prompt, for example "Instrumental, no vocals." A non-empty
negative_promptreturns HTTP 400 withpublic_reason=negative_prompt_unsupported. - Image:
image_urlis optional and must be a public HTTPS image.
What a good brief contains
The music brief template in the docs uses seven axes: emotion, genre or lineage, tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and era or production. The docs call them creative directions, not guaranteed output values, so listen to the result.
A seven-axis brief with timed sections is usually 400 to 900 characters. The script below counts a draft and flags the two things the API refuses, so you catch them before you submit. The limit is checked on the character count of the prompt; the counter uses Python's len, which matches for ordinary text but counts some emoji differently from the API.
prompt = (
"Hushed, slightly melancholic neo-soul, 72 BPM, D minor. "
"Rhodes through tape wow, soft sub bass, brushed snare, muted trumpet. "
"[0:00-0:12] Sparse Rhodes only. [0:12-0:30] Trumpet answers, bass enters. "
"A 30-second track. Instrumental, no vocals."
)
payload = {"model": "sume/music-auto", "prompt": prompt}
problems = []
if not 1 <= len(prompt) <= 5000:
problems.append(f"prompt length {len(prompt)} is outside 1-5000")
for banned in ("duration", "duration_seconds"):
if banned in payload:
problems.append(f"remove {banned}: write the length in the prompt")
if payload.get("negative_prompt"):
problems.append("remove negative_prompt: write exclusions positively")
print(len(prompt), "characters;", problems or "no problems found")Which model id
The router resolves sume/music-auto to an engine, Lyria 3.5 today, or you can pin lyria-3.5 or lyria-3-pro from GET /v1/music-router/models. An unknown model fails with 400 model_not_found and a catalog_url. Every request resolves through the router now, including those sent to the older Music 1.0 route.
After a rejection
If a prompt is rejected for policy, change the flagged content and keep the musical brief; the docs advise against flattening it into a generic bed. Retry only inside the budget you authorized, since each generation is billed.
What fits in 5,000 characters
Most music briefs are far shorter than the limit. A brief in the docs example style, with genre, tempo, key, instruments, one named moment and a no-vocals line, is a few hundred characters. The 5,000 limit starts to matter when you write a timed structure with a line per section, such as an intro, a build and an ending, each with its own time range.
Timed sections are the supported way to control length, since duration is rejected. A 2-minute track can be written as a 2-minute request in the prompt, or as time ranges like [0:00-0:30] Intro. Eight such lines at 150 characters each use 1,200 characters, well inside the limit.
Because the limit is a character count, count what you send, spaces and punctuation included, before you submit rather than after a rejection.
- Check the length in code:
len(prompt)must be between 1 and 5000. - Put exclusions in the positive prompt, since
negative_promptis rejected. - Long is not better; add a line only if it changes what you hear.
Sources
Related posts
More in Models
- MAI-Transcribe-2-Streaming 0.13 s to final: when the clock starts
The 0.13 second figure is measured from end of speech found by a VAD, and partials arrive in about 100 ms. Why neither number is the wait for a file transcript.
- MAI-Transcribe-2-Streaming 2.5% WER: what the test audio mix is
The 2.5% word error rate comes from a chunked-streaming index: 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. Here is how that maps to a recorded call.
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
Written by Sume