Add an intro and outro to every clip: audio concat plus a render
Add the same intro and outro to a batch of clips with Sume: concat the audio with timeline audio, then render three video slots using the returned offsets.

To add an intro and an outro to a clip with Sume, build the audio track first with POST /v1/timeline-1.0/audio (operation: concat), then render a Timeline 1.0 program with three video slots whose start values come from the segments[] offsets that call returns. The concat step does the arithmetic for you, so the video slots stay in step with the sound.
Why audio first
Timeline 1.0 is driven by its audio: audio.duration_seconds sets the length, and every video slot sits at a start on that clock. If you concat the intro sound, the clip sound and the outro sound into one file, the result lists each part's index, start and duration_seconds. Those are the numbers to put into the video slots.
| Item | Value |
|---|---|
| Concat parts per call | up to 20 |
| Audio concat price | $0.01 flat per call |
| Container | wav is sample-exact; mp3 re-adds padding |
| Render price | $0.10 per ceil output minute |
| Render audio length | audio.duration_seconds 1-1800 |
| Source URLs | Must be your workspace's media.sume.com artifacts or assets |
Steps
1. Detach the clip's sound with POST /v1/audio-detach ($0.01 a job; wav is the default). Do this once for the clip; the intro and outro sound can be reused across the whole batch. 2. Concat intro, clip and outro audio. A part may carry source_in and duration to take only a slice. 3. Read segments[] from the result. 4. Render with three slots: intro at the first offset, clip at the second, outro at the third, each duration equal to the segment's duration_seconds.
If you only need the join inside this one render, skip the concat job and pass the same files as audio.parts[] on the render itself; use the concat job when you want a reusable audio file. Because the render is a new artifact and the clip itself is untouched, you can reuse the same clip with a different intro without redoing anything.
{
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/joined.wav", "duration_seconds": 36},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 3},
{"source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 3, "duration": 30,
"transition": {"type": "fade", "duration": 0.3}},
{"source_url": "https://media.sume.com/artifacts/artf_demo/outro.mp4", "start": 33, "duration": 3,
"transition": {"type": "fade", "duration": 0.3}}
]
}Batch tip
The intro and outro files never change, so detach their sound once and reuse the audio URLs. For each clip only the middle part of the concat differs. If every clip has the same length, you can reuse one plan and change only the source URL. Run a plan on the first clip, confirm billable_minutes, then submit the batch.
Limits
A slot is at least 0.2 s and the first slot must start at 0. A transition is at most 1 s. Run POST /v1/timeline-1.0/plan (unbilled) before the real render to catch a bad program. If a source is shorter than its slot the renderer pads or loops it and returns a soft warning, which a plan cannot predict, so read warnings[] on the result. An mp3 concat adds encoder padding and shifts offsets slightly; stay on wav for tight sync.
Sources
Related posts
More in Media tools
- Talking photo audio too large? The 10 MB Sume Fabric limit
Sume's veed/fabric-1.0 needs Sume-hosted audio under 10 MiB and up to 300 seconds. Size arithmetic for WAV and MP3, and how to split or compress.
- Which AI video model makes 1:1 square clips on Sume?
Kling 3.0, Wan 3.0, MiniMax H3, H3 Max and Grok Imagine list 1:1 in Sume's Videos panel; Auto does not. What to pin for square feed video.
- Captions with sound cues for deaf viewers: W3C checklist on Sume
W3C says captions carry speech and non-speech sound. Sume's STT burn covers speech only, so here is how to author the full cue list and burn it.
- caption_no_speech on a silent clip: burn Halloween text with cues
A silent AI clip fails video captions with caption_no_speech. Pass cues with text, start and end instead: a worked giveaway announcement on Sume at $0.20.
Written by Sume