MMAudio video-to-sound: CC-BY-NC weights and commercial use
MMAudio makes synchronized sound for video, but its checkpoints are CC-BY-NC 4.0 and default to 8 seconds. What that means for ads, and the Sume route.

MMAudio generates sound that matches a video, but you should not put its output into a paid ad on the strength of the code license. The repository says the code is MIT while "the checkpoints are released on Hugging Face under the CC-BY-NC 4.0 license", and the authors add: "We do not guarantee that the pre-trained models are suitable for commercial use. Please use them at your own risk." For client work, plan on a different sound source.
What does MMAudio do, and what are its limits?
The README describes a model that makes "synchronized audio given video and/or text inputs". It lists video-to-audio, text-to-audio and an experimental image-to-audio mode. One number to plan around is length: the default output and training duration is 8 seconds, and longer or shorter lengths may work but a large deviation "may result in a lower quality".
| Item | README says |
|---|---|
| Inputs | Video and/or text; image-to-audio is experimental |
| Code license | MIT |
| Checkpoint license | CC-BY-NC 4.0 on Hugging Face |
| Default output length | 8 seconds |
| Commercial suitability | Not guaranteed; use at your own risk |
Why does CC-BY-NC matter for a social ad?
Non-commercial licenses cover research, study and personal projects. An ad exists to sell something, so the usual advice is to avoid NC weights for it. We are not lawyers and the README itself declines to promise commercial suitability, which is reason enough to pick another source for client deliverables.
What can Sume do for sound on a video?
The Sume docs we read describe no standalone sound-effects model. What they describe is video with built-in sound. Video Router says gemini-omni-flash-1.1 runs 3 to 10 seconds with native synced audio always on, and that generate_audio: false is rejected for it. The same row has a video_to_video edit mode that takes a video_url and a prompt describing the edit.
Once a clip exists you can pull its sound out. Audio detach turns the track of one Sume-hosted video into a wav by default, at $0.01 per job, so you can audition the sound or reuse it as the Timeline spine. Check that a clip has sound first with video inspect, where probe.has_audio answers it.
Which route fits which job?
Match the job to the cheapest honest tool:
- A sound to fit a clip you already own, for a personal project: MMAudio is a fair experiment under its NC terms.
- A clip with sound made from scratch for an ad: a video model with native audio, then detach if you need the track alone.
- A bed under a voice: a Music Router job and the Timeline
soundtrackfield. - A named, real-world sound: a recorded library with a license you can show.
What should you keep on file?
Keep the model name, the license text you read and the date. For Sume jobs, the job record shows which model produced the clip. A short note next to each deliverable saves a long email later.
What does 8 seconds mean for a real edit?
Most ads have several cuts and a runtime of 15 to 30 seconds, so an 8-second default means you generate sound per shot and stitch it. The README warns that large deviations from the training length may lower quality, which argues for per-shot generation instead of one long pass. Stitching sound across hard cuts is where gaps and level jumps appear, so plan for a mixing step either way.
Sume's side of that is the Timeline: when audio is already Sume-hosted, Timeline audio can concat up to 20 parts into one gapless file at a flat $0.01 per job, and the Timeline render takes it as audio.url.
Sources
Related posts
More in Models
- Retirement notice: Anthropic 60 days, OpenAI 6 months. Pin the model?
Anthropic promises at least 60 days' notice and OpenAI at least 6 months for GA models. Whether to pin or omit the model field in a saved Sume Format.
- Which Sume image model lists the most aspect ratios?
Ideogram V3 and 4.5 list 15 ratios; Nano Banana 2 lists 14 plus auto, including 8:1 and 1:8. Ratio counts for all 19 Sume image rows, with what to pick.
- Music Router routed_model: log which engine made each track
With sume/music-auto, job.request.routed_model names the engine that ran, such as lyria-3.5. Why to store it per track and how to read it from a job.
- Nano Banana Pro interleaved text and images vs Sume image output
Google documents Nano Banana Pro returning text blocks with illustrations in one answer. Sume's Image API documents image results only; here is the workaround.
Written by Sume