Replace an ad read in a podcast episode: one concat job, no editor
Swap a 30-second ad in a finished podcast file by concatenating three parts with source_in and duration on Sume Timeline audio. $0.01 flat, wav or mp3.

To replace an ad read in a finished podcast episode, run one Timeline audio concat job with three parts: the episode up to the ad, the new ad, and the episode from the end of the old ad. Each part is { url, source_in, duration }, so the episode file is used twice and never needs to be cut into pieces first. The job costs a flat $0.01.
This works on audio files already on media.sume.com. Sume does not find the ad for you, so you supply the timestamps.
What does the request look like?
Say the old ad runs from 612.0 to 642.0 seconds. Part one takes the episode from 0 for 612 seconds. Part two is the replacement read. Part three takes the episode from source_in 642 with no duration, which means the rest of the file. The parts must share one channel layout, so a mono ad into a stereo episode fails with audio_parts_channel_mismatch until you match them.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ep-212-ad-swap-v1" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/ep212.wav", "source_in": 0, "duration": 612 },
{ "url": "https://media.sume.com/artifacts/artf_demo/new-ad.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/ep212.wav", "source_in": 642 }
],
"output": { "format": "wav" }
}'The result is kind: timeline_audio with one audio_url, a total duration_seconds, and segments[] telling you where each part landed. The second segment's start is the new ad's position, which is useful if you publish chapter markers or ad-break timestamps.
Should the output be wav or mp3?
The default is wav, sample-exact pcm_s16le. mp3 is smaller, but it re-adds priming padding at every edge it joins. With three parts that is two seams, so a final delivery file can be mp3 if your host wants a small upload. If you will join or edit the file again, keep wav and encode once at the very end.
Apple and Spotify spec details are covered in Podcast loudness: Sume does not normalize loudness, so a replacement ad recorded louder than the episode will stay louder.
What are the hard limits?
| Limit | Value | What happens past it |
|---|---|---|
| Parts per job | 1 to 20 | Request refused |
| Produced audio | 1800 seconds (30 minutes) | Job fails with output_duration_exceeded |
| Source location | This workspace's media.sume.com audio | Off-host URLs are rejected; import first with POST /v1/media-imports |
| Price | $0.01 flat per job | Confirm in GET /v1/catalog |
The 30-minute ceiling is the one that bites. A 45-minute episode cannot be produced in one job. Split the episode into two halves with operation: split, apply the swap to the half that holds the ad, and concat the halves back only if the total fits; otherwise publish the result as two parts or do the final join in your own editor.
What if the episode is a video file?
Run Audio detach first to get the audio track as a durable wav or mp3 (a flat $0.01). Detach output is capped at 900 seconds per job, so a long video needs ranges. After the swap you have an audio file only; putting it back under the picture is a Timeline render job with the new file as the audio spine.
Everything runs asynchronously by default. Pass mode: sync to wait up to 30 seconds for a finished job, or poll as in Jobs and results.
Sources
Related posts
More in Use cases
- Synthesia Sessions alternative: avatar briefing clips by API
Synthesia launched Sessions for live avatar roleplay. Sume renders scripted avatar clips instead: what that covers in sales training, and what it does not.
- SynthID Detector: how to check a video for a watermark today
Google DeepMind's SynthID page says a detector portal exists with an early-tester waitlist, and Gemini can check uploads. What that means for a video team.
- Tennessee ELVIS Act and AI voice: a checklist for TTS projects
Tennessee's ELVIS Act, signed 21 March 2024, extends protection to voice. A consent checklist for AI voiceover, using the governor's page and Sume's TTS docs.
- Text to speech from a PDF with an API: extract, chunk, narrate
Sume's TTS takes text, not PDF files. Extract the text, split it under 20,000 characters, narrate each chunk, then join the audio. Costs and limits included.
Written by Sume