Mirelo SFX video-to-sound vs how Sume adds sound to video
Mirelo generates sound effects synced to existing video. Sume has no video-to-SFX route; here is what it does offer for sound on a clip, and where each fits.

Mirelo generates sound effects for footage you already have; Sume does not. Sume has no video-to-sound route, so for an existing clip you either generate a music bed, record or synthesize a voiceover, or generate a new clip on a model that makes its own audio.
Mirelo's home page was read on 2026-10-02. Sume facts come from the docs linked in the sources.
What does Mirelo offer?
Mirelo's page describes a platform that analyzes video scenes and generates sound effects and music synced to what is on screen. Its main tool is Mirelo SFX, with Studio as a web editor that can generate SFX from footage, extend existing effects, create ambient soundscapes from text and remove audio artifacts such as clicks, hisses and background noise.
Integrations listed are plugins for Adobe Premiere Pro, DaVinci Resolve and Roblox Studio, an MCP for Claude, ChatGPT and Cursor, and a REST API. The page lists a free tier in Studio and points to separate pricing pages that we did not read, so no price is quoted here. Announcements it lists for July 2026 are the MCP launch, Gibberize and Audio-to-MIDI.
Does Sume generate sound effects for a video?
No. There is no Sume endpoint that watches a clip and returns matching foley. What Sume ships for sound is below.
Native audio on generated video. In the Video Router, gemini-omni-flash-1.1 returns 3 to 10 second clips with native synced audio always on, and generate_audio: false is rejected. The Auto family also always generates audio, per the Video 1.0 docs. This is sound for a new clip, not for footage you shot.
Music from a prompt. Music 1.0 and the Music Router generate a track from text with an optional image_url, at a fixed $0.125 per accepted generation.
Mixing. Timeline 1.0 accepts a soundtrack bed with gain_db, loop, fade_out_seconds up to 10 and duck_db 0 to 20 under a voiceover spine.
How do the two approaches compare?
The table lines up the jobs you might be doing.
| You have | Mirelo (per its page) | Sume |
|---|---|---|
| A finished clip needing footsteps, doors, ambience | SFX generated from the video | No route; add a music bed or voiceover instead |
| A text description of an ambience | Ambient soundscapes from text in Studio | Describe the ambience in a music prompt and check the result |
| A noisy recording to clean | Artifact removal in Studio | No denoise route |
| A new clip with sound | Not described on the page | Video Router models with native audio |
| Music under a voiceover | Music listed on the page | Music job plus a timeline soundtrack with ducking |
How would I add a music bed to an existing clip on Sume?
Pull a still with video frames or read the probe from video inspect, write a music brief that names tempo, instruments and the one moment where the track should change, and generate it. Music accepts no duration field, so ask for the length in the prompt and trim in the timeline.
Then render with a soundtrack entry. Sume's render keeps ffmpeg server-side, so you send fields, not filter graphs.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: bed-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Light foley-forward comedy bed, 104 BPM, C major. Pizzicato strings and a woodblock. A 20-second track. Instrumental, no vocals.",
"mode": "async"
}'Which should I choose?
If the job is matching effects to existing footage, Mirelo is built for it and Sume is not. If you generate the clips too, Sume's native-audio models and the timeline keep picture, voice and music in one billing and one job system. Many teams will use both: generated or shot footage, effects from a specialist and the final mix in a render.
Two honest caveats. First, Mirelo's page says nothing about price or limits in the part we read, so cost cannot be compared yet. Second, Sume's native audio is generated with the picture and cannot be regenerated alone: if you dislike the sound on a clip you regenerate the clip, or you replace the audio spine in a render with a voiceover or music file you supply. For a clip that already works visually, that replacement path is usually cheaper than a new generation, because a timeline render is billed at $0.10 per output minute rather than per generation.
Sources
Related posts
More in Comparisons
- Mubert Render 25-minute tracks and licence limits vs Sume Music
Mubert Render makes tracks up to 25 minutes on paid plans but bars streaming release. Sume Music makes song-length tracks at $0.125 each; loop for longer.
- Murf API alternative for video voiceover: where Sume fits
Murf is a voice API with Falcon, dubbing and 150+ voices. Sume has an async TTS route and turns a script into a talking video; it has no dubbing API.
- Mux thumbnail time and fit_mode vs Sume video frames at and max_edge
Mux gets a thumbnail from a playback ID URL with time, width and fit_mode. Sume video-frames takes at[] times or fps and returns durable image files.
- Nano Banana 2: five characters, 14 objects vs Sume reference slots
Google says Nano Banana 2 keeps up to five characters and 14 objects consistent. Sume caps reference images per model; see what you can send and where it stops.
Written by Sume