How to mix voice with background music into one file

Mix voice with background music: set the music below the voice, duck it under speech, fade it out, export one file. The Sume way: render, then detach.

6 min readSume
All posts

To mix a voice with background music, put the music on a second track well below the voice, dip it further while the voice speaks (ducking), fade it out at the end, and export both tracks as one audio file. With Sume's API the mixer is a Timeline 1.0 render: the voice is the audio spine, the music is the soundtrack, and one Sume-hosted still fills the picture. Audio detach then pulls the mixed track out of the MP4 as a WAV or MP3.

The facts come from the Timeline 1.0, Audio detach, and Timeline audio docs and the Timeline field descriptions in the Sume API reference, read on 2026-09-28; anything called current behavior is read from Sume's code. What each music setting does to a video's sound is covered in add background music to a video; this post ends with an audio file.

Can I add background music to audio I already recorded?

In any multitrack audio editor, yes: voice on one track, music on another, the music lowered, dipped under speech and faded, then one export.

Sume's route works on files that are already on Sume: every URL in a render must be your workspace's media.sume.com artifact or asset, such as an earlier Sume job's output, and Sume mirrors generated outputs to its own media URLs. That fits a voice made with text to speech, as in how to make an AI voiceover, and a bed from Sume's music generation, whose prompt guide says to close the brief with “Instrumental, no vocals.” and add “no spoken word” under narration.

How loud should background music be under a voice?

Low enough that every word stays clear; how low is a judgment call for each mix, so listen on the kind of speakers or earbuds your audience will use. Sume's default bed level is −16 dB, which the API reference calls a bed level under a spine, and duck_db (0–20) dips the bed further while the voice is speaking. Add background music to a video explains the duck and the bed's other settings, loop and fade_out_seconds.

How do I mix them with the Sume API?

Send one render with the voice as audio.url, the music as soundtrack, and one still in video[]:

  • Set audio.duration_seconds to the voice's length; the output always runs exactly that long. A TTS result requested with word timestamps reports the length as duration_seconds in current code.
  • video[] needs at least one slot, and a still is a static hold, so one Sume-hosted image from 0 to the end carries the picture.
  • The mix inherits the voice file's sample rate and channel count, so send the voice at the quality you want to keep.
  • The render runs as a job: poll GET /v1/jobs/:id/status, then read the MP4's video_url from GET /v1/jobs/:id/result.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: intro-mix-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
      "duration_seconds": 42
    },
    "soundtrack": {
      "url": "https://media.sume.com/artifacts/artf_demo/music.mp3",
      "loop": true,
      "duck_db": 8,
      "fade_out_seconds": 4
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/cover.png", "start": 0, "duration": 42 }
    ]
  }'

How do I get the mix as an MP3 or WAV file?

Send the render's video_url to POST /v1/audio-detach. It returns a new audio file and leaves the video untouched. Keep WAV if the mix will be joined again, as for an intro below; MP3 is the smaller file. WAV vs MP3 compares the two.

Audio detach options for the finished mix, from Audio detach, read 2026-09-28.
FieldValuesOmitted
formatwav (pcm_s16le) or mp3 (128 kbps)wav
sample_rate16000, 44100, or 48000The source's rate
channelssource or monosource
range{ start, end? } in secondsThe whole track
Length capsSource up to 1,800 s; output up to 900 s per jobA longer mix needs a range per job
curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: intro-mix-mp3-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/mix.mp4",
    "format": "mp3",
    "sample_rate": 44100
  }'

Can the music play before the voice starts or after it ends?

Not from the soundtrack alone. It has no start-time field, so it starts with the output, and its fade runs over the last seconds of the spine. For a music-only intro or outro, as in a podcast opening, make a short sting, as in AI jingle generator, then join sting, mix, and sting in order with a Timeline audio concat: 1 to 20 Sume-hosted parts, with no silence added at the seams. The parts must share one channel layout, or the join fails with audio_parts_channel_mismatch.

What does mixing voice and music cost?

The render's public rate is $0.10 per output minute on API pricing, reserved as ceil(audio.duration_seconds / 60) minutes, and each detach is $0.01 per job, the rate on its docs page, which says to confirm it in GET /v1/catalog. Making the voice and the music, and any concat, are separate jobs at their own rates, and every job is billed plus a 5.5% agent fee by default.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume