Sonilo segment-level music controls vs Sume section markers
Sonilo's text-to-music lets you set styles and moods per section. On Sume you write section markers like [0:00-0:30] Intro: inside one 5000-character prompt.

Sonilo's June 22, 2026 release says its text-to-music model, also on fal.ai, has segment-level controls that let you define different musical styles, moods and structures across sections of a track. Sume has no separate per-segment field. You describe the sections inside one prompt, using timestamped markers such as [0:00-0:30] Intro: ..., which the Music 1.0 docs give as the way to steer structure and length. The Sonilo statement is from its press release; the release does not list the field names.
How do section markers work on Sume?
They are text in the prompt, not parameters. The Music Router docs say to steer length in the prompt ("a 2-minute track" or markers like [0:00-0:30] Intro: ...), and that exclusions go in the positive prompt because a non-empty negative_prompt returns 400. The prompt can be 1 to 5000 characters, so a few sections fit comfortably.
| Question | Sonilo release | Sume docs |
|---|---|---|
| How sections are set | Segment-level controls | Timestamped markers inside prompt |
| Length control | Not stated for text model | Steer in the prompt; duration fields rejected |
| Prompt size | Not stated | 1 to 5000 characters |
| Guarantee | Not stated | Markers are creative directions; verify the audio |
What does a sectioned prompt look like?
Keep each section to one mood and one instrument change, and name one arc moment, so the model has something concrete. The Music 1.0 docs also suggest a tempo as a number and a key.
[0:00-0:10] Intro: sparse Rhodes, hushed, 84 BPM, D minor.
[0:10-0:25] Build: brushed drums enter, muted trumpet answers the Rhodes.
[0:25-0:30] Resolve: drums drop out, one held chord.
A 30-second track. Instrumental, no vocals.What if the sections come out wrong?
Sume's docs call prompt directions creative guidance and tell you to verify the generated audio. There is no seed, so rerunning the same prompt can give a different result. Revise the section text and run a new job; the stored guide AI song structure tags covers the tag styles that Lyria accepts.
Sources
Related posts
More in Comparisons
- Sonilo video-to-music on fal.ai: 600 s of footage vs Sume's route
Sonilo's video-to-music model scores footage up to 600 seconds on fal.ai. Sume has no video-to-music call: inspect the clip, write a prompt, mix with Timeline.
- Soundraw Artist Pro WAV and stems vs Sume Music at $0.125 a track
Soundraw offers WAV and stems from Artist Pro and API access only on Enterprise. Sume Music returns one audio file per call at $0.125 and no stems. Compared.
- Speech-to-text price per audio hour: xAI, OpenAI, Sume
Per audio hour xAI lists $0.10, OpenAI mini $0.18, OpenAI 4o $0.36 and Sume video_inspect transcripts $0.60. Dated 2026-10-01, with honest limits.
- Speechify API vs Sume: speech marks and streaming vs video jobs
Speechify's API streams speech with word-level marks and voice cloning. Sume documents no standalone speech endpoint, but accepts audio for talking-video jobs.
Written by Sume