Sonilo video-to-music on fal.ai: 600 s of footage vs Sume's route
Sonilo's video-to-music model scores footage up to 600 seconds on fal.ai. Sume has no video-to-music call: inspect the clip, write a prompt, mix with Timeline.

Sonilo announced on June 22, 2026 that its video-to-music model is live on fal.ai, and says it analyzes a video's pacing, motion and emotional arc and generates music matched to the video's exact duration, for footage up to 600 seconds. Sume has no video-to-music call. On Sume you describe the music yourself in a text prompt on the Music Router, optionally after probing the clip with video inspect, then place the track with a Timeline soundtrack. The Sonilo facts are from its press release.
What does the Sonilo release say?
It says each generation produces several soundtrack options for the same footage, the music is delivered as a separate audio track so volume can be changed without touching dialogue or effects, and original speech can be preserved. It also gives internal figures (an 87% first-attempt acceptance and a 16% engagement increase); those are the vendor's own, not independent measurements, and are not used here for any comparison.
| Question | Sonilo release | Sume |
|---|---|---|
| Input | Video footage, or text on its text model | Text prompt plus optional image_url |
| Length match | Composed to the video's exact duration | No duration field; steer length in the prompt, then trim in Timeline |
| Max footage | Up to 600 seconds on fal.ai | Not applicable |
| Output | Separate audio track | Audio artifact, mixed under a spine with soundtrack |
| Options per run | Multiple soundtrack options | One audio artifact per accepted job |
How do I score a clip on Sume?
Probe the clip first. Video inspect reads one media.sume.com clip and returns probe facts, stills and optional transcription; it does not describe scenes semantically, so you read the stills and decide the mood. Write a brief with tempo, key, instruments and an arc moment, ask for a track about as long as the clip, then render with soundtrack (gain_db, fade_out_seconds, duck_db).
What about getting several options?
Submit more than one job with different briefs and an Idempotency-Key per intended job; each accepted job is a separate generation at the fixed Music price. Pick the one that fits the cut. This is a manual version of Sonilo's multiple options, and it is yours to run and review.
Does Sume say anything about licensing?
The Sonilo release says its models are trained on professionally licensed content and outputs are available for commercial use under its terms. The Sume docs pages used here do not make an equivalent statement, so check your own platform and rights requirements before publishing.
Sources
Related posts
More in Comparisons
- Soundraw Artist Pro WAV and stems vs Sume Music at $0.125 a track
Soundraw offers WAV and stems from Artist Pro and API access only on Enterprise. Sume Music returns one audio file per call at $0.125 and no stems. Compared.
- Speech-to-text price per audio hour: xAI, OpenAI, Sume
Per audio hour xAI lists $0.10, OpenAI mini $0.18, OpenAI 4o $0.36 and Sume video_inspect transcripts $0.60. Dated 2026-10-01, with honest limits.
- Speechify API vs Sume: speech marks and streaming vs video jobs
Speechify's API streams speech with word-level marks and voice cloning. Sume documents no standalone speech endpoint, but accepts audio for talking-video jobs.
- Stable Audio DAW plugin: BPM sync and takes vs Sume Music
Stability's Stable Audio plugin runs in Logic Pro and Ableton Live with BPM sync. Sume has no plugin or tempo field; Music Router returns an audio file by API.
Written by Sume