MMAudio video-to-sound: CC-BY-NC weights and commercial use

MMAudio makes synchronized sound for video, but its checkpoints are CC-BY-NC 4.0 and default to 8 seconds. What that means for ads, and the Sume route.

4 min readSume
All posts

MMAudio generates sound that matches a video, but you should not put its output into a paid ad on the strength of the code license. The repository says the code is MIT while "the checkpoints are released on Hugging Face under the CC-BY-NC 4.0 license", and the authors add: "We do not guarantee that the pre-trained models are suitable for commercial use. Please use them at your own risk." For client work, plan on a different sound source.

What does MMAudio do, and what are its limits?

The README describes a model that makes "synchronized audio given video and/or text inputs". It lists video-to-audio, text-to-audio and an experimental image-to-audio mode. One number to plan around is length: the default output and training duration is 8 seconds, and longer or shorter lengths may work but a large deviation "may result in a lower quality".

MMAudio facts from the project README, https://github.com/hkchengrex/MMAudio, read 2026-10-03.
ItemREADME says
InputsVideo and/or text; image-to-audio is experimental
Code licenseMIT
Checkpoint licenseCC-BY-NC 4.0 on Hugging Face
Default output length8 seconds
Commercial suitabilityNot guaranteed; use at your own risk

Why does CC-BY-NC matter for a social ad?

Non-commercial licenses cover research, study and personal projects. An ad exists to sell something, so the usual advice is to avoid NC weights for it. We are not lawyers and the README itself declines to promise commercial suitability, which is reason enough to pick another source for client deliverables.

What can Sume do for sound on a video?

The Sume docs we read describe no standalone sound-effects model. What they describe is video with built-in sound. Video Router says gemini-omni-flash-1.1 runs 3 to 10 seconds with native synced audio always on, and that generate_audio: false is rejected for it. The same row has a video_to_video edit mode that takes a video_url and a prompt describing the edit.

Once a clip exists you can pull its sound out. Audio detach turns the track of one Sume-hosted video into a wav by default, at $0.01 per job, so you can audition the sound or reuse it as the Timeline spine. Check that a clip has sound first with video inspect, where probe.has_audio answers it.

Which route fits which job?

Match the job to the cheapest honest tool:

  • A sound to fit a clip you already own, for a personal project: MMAudio is a fair experiment under its NC terms.
  • A clip with sound made from scratch for an ad: a video model with native audio, then detach if you need the track alone.
  • A bed under a voice: a Music Router job and the Timeline soundtrack field.
  • A named, real-world sound: a recorded library with a license you can show.

What should you keep on file?

Keep the model name, the license text you read and the date. For Sume jobs, the job record shows which model produced the clip. A short note next to each deliverable saves a long email later.

What does 8 seconds mean for a real edit?

Most ads have several cuts and a runtime of 15 to 30 seconds, so an 8-second default means you generate sound per shot and stitch it. The README warns that large deviations from the training length may lower quality, which argues for per-shot generation instead of one long pass. Stitching sound across hard cuts is where gaps and level jumps appear, so plan for a mixing step either way.

Sume's side of that is the Timeline: when audio is already Sume-hosted, Timeline audio can concat up to 20 parts into one gapless file at a flat $0.01 per job, and the Timeline render takes it as audio.url.

Sources

Related posts

More in Models

All Models posts

Written by Sume