Replace audio from 12.4 s to 14 s of a clip with Timeline parts
Seedance 2.5 lists timestamp-level audio editing. On Sume you can detach a clip's audio and rebuild it from parts with a replaced line in the middle.

Seedance 2.5 lists timestamp-level editing of audio and video (ByteDance Seed, read 2026-10-05). If your goal is narrower, replacing the audio between 12.4 and 14 seconds of a finished clip, you do not need a video model for it. On Sume you detach the clip's audio, then build a new audio spine from three parts (before, replacement, after) and render the clip over it with Timeline 1.0. The ffmpeg-only media routes involve no provider inference.
The four steps
All URLs must already be on media.sume.com for your workspace. Import files first with POST /v1/media-imports, and send an Idempotency-Key on every write.
- Detach the clip's audio to a wav with
POST /v1/audio-detach(default sample-exact wav). Public rate: $0.01 per job. - Produce the replacement line (for example with a text-to-speech model) as a Sume-hosted wav about 1.6 seconds long.
- Render with Timeline 1.0 using
audio.parts[]: slice 0 to 12.4 from the original, then the replacement, then the original from 14 on. - Listen at the two seams. Wav parts join in the sample domain with no gap, so any click comes from the content, not the join.
The render call
audio.parts[] takes up to 20 gapless slices, each with a url and optional source_in and duration. The video slot uses the original clip for the full 30 seconds. The three part lengths must add up to at least audio.duration_seconds, or the render is refused with audio_parts_shorter_than_duration.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audio-swap-001" \
-d '{
"audio": {
"duration_seconds": 30,
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/orig.wav", "source_in": 0, "duration": 12.4 },
{ "url": "https://media.sume.com/artifacts/artf_demo/new-line.wav", "duration": 1.6 },
{ "url": "https://media.sume.com/artifacts/artf_demo/orig.wav", "source_in": 14, "duration": 16 }
]
},
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 0, "duration": 30 }
]
}'Pitfalls with parts and seams
Three things go wrong most often. First, a replacement line that is longer or shorter than the gap shifts everything after it, so measure the replacement and adjust the third part's source_in to match. Second, a stereo original joined with a mono replacement is refused for channel layout; the audio join docs list audio_parts_channel_mismatch. Detach both as the same channel layout, for example channels: "mono". Third, mp3 adds priming padding at each edge, so keep wav for anything you will join.
- Match channel layouts across parts.
- Keep wav for joins.
- Measure the replacement; the spine length is authoritative.
Why not regenerate the whole clip
A new generation produces a new take, with new motion and new faces. If the picture is approved and only a word is wrong, this audio route keeps every pixel as it was. It also costs about a dime in ffmpeg jobs rather than another model run. Use a regeneration only when the picture has to change. If the replacement line is a different voice from the original, expect a noticeable change at the seams; a short music bed under the whole clip, added as the soundtrack with duck_db, can smooth it. That field needs a real spine, which parts provide.
Where the model route still helps
If the picture must change in that window, say a new product shown at 12.4 s, audio parts are not enough. Cut the clip around the window with video trim (start plus end or duration, $0.02 per job), generate a replacement shot, and join the three clips as video[] slots. The Sume docs recommend putting the trim output into video[] with source_in 0.
Keep one rule: match the replacement's length to the gap, because Timeline coverage is declared and the spine length is authoritative.
Sources
Related posts
More in Use cases
- Swap a product's colorway in a finished clip: Gemini Omni edit
One approved ad, five colorways: gemini-omni-flash-1.1 on Sume edits a source video from a prompt. Input rules, what it cannot carry, and a batch plan.
- SynthID not detected in Gemini: that does not mean the video is real
Gemini's checker only recognizes Google AI content, so no watermark means no Google mark was found, not that a person filmed it. Know its limits.
- Talking-head ad: Omni audio, Kling audio on, or Avatar Video?
Pick by length and face. Avatar Video runs 4 to 60 s from a script; Omni makes 3 to 10 s with native audio; Kling 3 adds audio at $0.21 a second.
- Talking photo: H3 Max lip sync, Kling motion control or Recast?
Lip sync follows your audio, motion control follows a reference video, Recast swaps a person. 10-second prices on Sume: $1.00, $1.58 and $3.75 at 768p.
Written by Sume