Join voiceover takes into one gapless wav: timeline audio concat

Timeline audio concat joins up to 20 hosted audio parts sample-exact into one reusable wav for $0.01 per job and returns segment offsets.

3 min readSume
All posts

Send operation: "concat" and up to 20 parts[] to POST /v1/timeline-1.0/audio and you get one gapless file back at $0.01 per job. The join is in the sample domain, with no re-synthesis and no silence at the seams (Timeline audio).

Request

Each part is {url, source_in?, duration?}. The second part below skips its first 0.1 s and keeps 1.8 s.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: vo-concat-001" \
  -d '{
    "operation": "concat",
    "parts": [
      {"url": "https://media.sume.com/artifacts/artf_demo/line1.wav"},
      {"url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8}
    ],
    "output": {"format": "wav"}
  }'

What comes back

The result is kind: timeline_audio with one audio_url, duration_seconds and segments[] holding index, start and duration_seconds for each part. Use those offsets to re-base video[].start in your timeline.

Concat or audio.parts?

Sume docs, read 2026-10-08
NeedUseCost
Reusable merged filetimeline-audio concat$0.01 per job
Join only inside one renderaudio.parts[] on Timeline renderIncluded in the render
Slice one file into rangestimeline-audio split$0.01 per job

Format notes

  • wav (pcm_s16le) is the default and sample-exact; mp3 is smaller and adds priming padding at every edge, so keep wav if you will join again.
  • All parts must share a channel layout: audio_parts_channel_mismatch.
  • The produced audio is at most 1800 s.

After the join

Pass the returned audio_url as audio.url on a Timeline 1.0 render, with audio.duration_seconds set to the returned duration_seconds. The same file also works as Avatar 1.0 image-to-video audio. If the render is the only place the voice is used, skip this job and send the takes as audio.parts[] instead.

Need the opposite? operation: "split" takes one top-level url and 1 to 20 ranges[], each {start, end?}, and returns a separate audio_url per segment. It is also $0.01 per job. If the source is a talking-head MP4, run audio detach once and split the audio it returns.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume