Replace audio from 12.4 s to 14 s of a clip with Timeline parts

Seedance 2.5 lists timestamp-level audio editing. On Sume you can detach a clip's audio and rebuild it from parts with a replaced line in the middle.

5 min readSume
All posts

Seedance 2.5 lists timestamp-level editing of audio and video (ByteDance Seed, read 2026-10-05). If your goal is narrower, replacing the audio between 12.4 and 14 seconds of a finished clip, you do not need a video model for it. On Sume you detach the clip's audio, then build a new audio spine from three parts (before, replacement, after) and render the clip over it with Timeline 1.0. The ffmpeg-only media routes involve no provider inference.

The four steps

All URLs must already be on media.sume.com for your workspace. Import files first with POST /v1/media-imports, and send an Idempotency-Key on every write.

  • Detach the clip's audio to a wav with POST /v1/audio-detach (default sample-exact wav). Public rate: $0.01 per job.
  • Produce the replacement line (for example with a text-to-speech model) as a Sume-hosted wav about 1.6 seconds long.
  • Render with Timeline 1.0 using audio.parts[]: slice 0 to 12.4 from the original, then the replacement, then the original from 14 on.
  • Listen at the two seams. Wav parts join in the sample domain with no gap, so any click comes from the content, not the join.

The render call

audio.parts[] takes up to 20 gapless slices, each with a url and optional source_in and duration. The video slot uses the original clip for the full 30 seconds. The three part lengths must add up to at least audio.duration_seconds, or the render is refused with audio_parts_shorter_than_duration.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-swap-001" \
  -d '{
    "audio": {
      "duration_seconds": 30,
      "parts": [
        { "url": "https://media.sume.com/artifacts/artf_demo/orig.wav", "source_in": 0, "duration": 12.4 },
        { "url": "https://media.sume.com/artifacts/artf_demo/new-line.wav", "duration": 1.6 },
        { "url": "https://media.sume.com/artifacts/artf_demo/orig.wav", "source_in": 14, "duration": 16 }
      ]
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 0, "duration": 30 }
    ]
  }'

Pitfalls with parts and seams

Three things go wrong most often. First, a replacement line that is longer or shorter than the gap shifts everything after it, so measure the replacement and adjust the third part's source_in to match. Second, a stereo original joined with a mono replacement is refused for channel layout; the audio join docs list audio_parts_channel_mismatch. Detach both as the same channel layout, for example channels: "mono". Third, mp3 adds priming padding at each edge, so keep wav for anything you will join.

  • Match channel layouts across parts.
  • Keep wav for joins.
  • Measure the replacement; the spine length is authoritative.

Why not regenerate the whole clip

A new generation produces a new take, with new motion and new faces. If the picture is approved and only a word is wrong, this audio route keeps every pixel as it was. It also costs about a dime in ffmpeg jobs rather than another model run. Use a regeneration only when the picture has to change. If the replacement line is a different voice from the original, expect a noticeable change at the seams; a short music bed under the whole clip, added as the soundtrack with duck_db, can smooth it. That field needs a real spine, which parts provide.

Where the model route still helps

If the picture must change in that window, say a new product shown at 12.4 s, audio parts are not enough. Cut the clip around the window with video trim (start plus end or duration, $0.02 per job), generate a replacement shot, and join the three clips as video[] slots. The Sume docs recommend putting the trim output into video[] with source_in 0.

Keep one rule: match the replacement's length to the gap, because Timeline coverage is declared and the spine length is authoritative.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume