Turn a music composition plan into a Sume time-range prompt (Python)
ElevenLabs music_v2_5 plans allow 6,132 characters in up to 30 lines. Sume's Music prompt takes 5000 characters. A Python converter for time ranges.

An ElevenLabs composition plan for music_v2_5 can run to 6,132 characters across up to 30 lines of 200 characters each. Sume's Music prompt is one text field of 1 to 5000 characters, so a plan has to be folded into time-range sections and checked against that cap before you submit.
The converter below takes a list of sections with durations and writes the [0:00-0:15] Name: text format that Sume's docs show.
The two limits
From the ElevenLabs changelog and the Sume Music docs, read 2026-10-03.
| Item | ElevenLabs music_v2_5 plan | Sume Music 1.0 prompt |
|---|---|---|
| Model id | music_v2_5 | Lyria through Sume Music |
| Total size | Up to 6,132 characters | 1 to 5000 characters |
| Structure | Up to 30 lines, 200 characters each | One prompt, time-range markers |
| Length control | Set in the plan | No duration field; state it in the prompt |
| Exclusions | Not covered here | In the positive prompt; non-empty negative_prompt returns 400 |
The converter
The input is a plain list you build from your plan: a name, a length in seconds and a line of direction per section. The script does not parse any vendor JSON, since this page has no field list for it. It accumulates start times, writes the markers, appends a closing clause, and refuses to emit a prompt over 5000 characters.
def fmt(t):
return f"{t // 60}:{t % 60:02d}"
def to_prompt(style, sections, total_label=None, closing="Instrumental, no vocals."):
t, lines = 0, []
for name, seconds, text in sections:
lines.append(f"[{fmt(t)}-{fmt(t + seconds)}] {name}: {text}")
t += seconds
head = f"{style} A {total_label or fmt(t)} track."
prompt = "\n".join([head, *lines, closing])
if len(prompt) > 5000:
raise ValueError(f"prompt is {len(prompt)} characters; the cap is 5000")
return prompt
sections = [
("Intro", 15, "sparse Rhodes and brushes, no trumpet."),
("Main", 20, "trumpet answers the Rhodes, bass thickens."),
("Outro", 10, "drums drop out, Rhodes alone."),
]
print(to_prompt("Warm lo-fi hip hop, 84 BPM, C minor.", sections, "45-second"))Fitting a long plan
A plan that is near the ElevenLabs maximum will not fit as it is.
- Merge neighbouring sections that share an instrument set. Fewer markers say the same thing.
- Move tempo, key and instrument list into the single style line, and keep per-section text to what changes.
- Drop text that is a label rather than a direction. A section called Verse 2 with no musical content adds length without information.
- Keep the exclusion clause. Sume's docs put it in the positive prompt.
Check the result before spending
Sume's docs describe a brief as a creative direction, not a guaranteed setting, and a finished job costs the fixed per-generation price whatever the prompt length. Listen to the first take for the section boundaries, then adjust the durations in the list and run the converter again rather than editing the prompt by hand, so the prompt and the plan stay in step.
Sources
Related posts
More in Media tools
- Cut a clip into 30-second chunks for 30-second post-processing caps
Runway's SDR-to-HDR model and the Magnific upscaler take 30 seconds at most. Cut any clip into 30-second pieces with Sume video trim at $0.02 per cut.
- Lyria 3.5 output: MP3 or WAV at 44.1 kHz, and what Sume returns
Google lists MP3 by default or WAV, 44.1 kHz stereo, with a SynthID watermark. Sume's music docs say the artifact is typically audio/mpeg. Check it in code.
- Seedance 2.5 secondary edit vs trim-and-regenerate on Sume
BytePlus describes timestamp-level edits to Seedance 2.5 clips. Sume has no such edit field, so here is the trim-and-regenerate route and its limits.
- Stability AI's Series B and label backers: what it means for audio
Stability released Stable Audio 3.0 on 5/20/26 and raised a Series B on 8/25/26 with EA, Sony, UMG and WMG. A dated timeline and what to verify for video work.
Written by Sume