Mirelo SFX video-to-sound vs how Sume adds sound to video

Mirelo generates sound effects synced to existing video. Sume has no video-to-SFX route; here is what it does offer for sound on a clip, and where each fits.

5 min readSume
All posts

Mirelo generates sound effects for footage you already have; Sume does not. Sume has no video-to-sound route, so for an existing clip you either generate a music bed, record or synthesize a voiceover, or generate a new clip on a model that makes its own audio.

Mirelo's home page was read on 2026-10-02. Sume facts come from the docs linked in the sources.

What does Mirelo offer?

Mirelo's page describes a platform that analyzes video scenes and generates sound effects and music synced to what is on screen. Its main tool is Mirelo SFX, with Studio as a web editor that can generate SFX from footage, extend existing effects, create ambient soundscapes from text and remove audio artifacts such as clicks, hisses and background noise.

Integrations listed are plugins for Adobe Premiere Pro, DaVinci Resolve and Roblox Studio, an MCP for Claude, ChatGPT and Cursor, and a REST API. The page lists a free tier in Studio and points to separate pricing pages that we did not read, so no price is quoted here. Announcements it lists for July 2026 are the MCP launch, Gibberize and Audio-to-MIDI.

Does Sume generate sound effects for a video?

No. There is no Sume endpoint that watches a clip and returns matching foley. What Sume ships for sound is below.

Native audio on generated video. In the Video Router, gemini-omni-flash-1.1 returns 3 to 10 second clips with native synced audio always on, and generate_audio: false is rejected. The Auto family also always generates audio, per the Video 1.0 docs. This is sound for a new clip, not for footage you shot.

Music from a prompt. Music 1.0 and the Music Router generate a track from text with an optional image_url, at a fixed $0.125 per accepted generation.

Mixing. Timeline 1.0 accepts a soundtrack bed with gain_db, loop, fade_out_seconds up to 10 and duck_db 0 to 20 under a voiceover spine.

How do the two approaches compare?

The table lines up the jobs you might be doing.

Sound for video, Mirelo and Sume, read 2026-10-02
You haveMirelo (per its page)Sume
A finished clip needing footsteps, doors, ambienceSFX generated from the videoNo route; add a music bed or voiceover instead
A text description of an ambienceAmbient soundscapes from text in StudioDescribe the ambience in a music prompt and check the result
A noisy recording to cleanArtifact removal in StudioNo denoise route
A new clip with soundNot described on the pageVideo Router models with native audio
Music under a voiceoverMusic listed on the pageMusic job plus a timeline soundtrack with ducking

How would I add a music bed to an existing clip on Sume?

Pull a still with video frames or read the probe from video inspect, write a music brief that names tempo, instruments and the one moment where the track should change, and generate it. Music accepts no duration field, so ask for the length in the prompt and trim in the timeline.

Then render with a soundtrack entry. Sume's render keeps ffmpeg server-side, so you send fields, not filter graphs.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: bed-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Light foley-forward comedy bed, 104 BPM, C major. Pizzicato strings and a woodblock. A 20-second track. Instrumental, no vocals.",
    "mode": "async"
  }'

Which should I choose?

If the job is matching effects to existing footage, Mirelo is built for it and Sume is not. If you generate the clips too, Sume's native-audio models and the timeline keep picture, voice and music in one billing and one job system. Many teams will use both: generated or shot footage, effects from a specialist and the final mix in a render.

Two honest caveats. First, Mirelo's page says nothing about price or limits in the part we read, so cost cannot be compared yet. Second, Sume's native audio is generated with the picture and cannot be regenerated alone: if you dislike the sound on a clip you regenerate the clip, or you replace the audio spine in a render with a voiceover or music file you supply. For a clip that already works visually, that replacement path is usually cheaper than a new generation, because a timeline render is billed at $0.10 per output minute rather than per generation.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume