Insert a sponsor read into a podcast: TTS plus one audio concat
Generate the ad read with Sume TTS, then join it into the episode at the second you choose with one Timeline audio concat: sample-exact, $0.01 per join.

To insert a sponsor read into an existing podcast episode, generate the read with Sume TTS as a WAV, then call POST /v1/timeline-1.0/audio with operation: "concat" and three parts: the episode up to the insert point, the read, and the episode from the insert point on. The join is sample-domain with no silence at the seams, and costs a flat $0.01. The script below does the whole thing. It needs the episode on media.sume.com, because Timeline audio refuses off-host URLs.
Voice work for ads is getting cheaper to start: the October tracker lists Microsoft's MAI-Voice-2.1 Flash at $15 per 1M characters (read 2026-10-06). A 330-character read is half a cent there. On Sume it is $0.0157 at $0.0475 per 1,000 (pricing), which the Sume price book rounds up to whole cents, so $0.02 on the invoice. The ad read is never the expensive part. Getting it into the episode cleanly is.
Generate and join
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
import urllib.error
EPISODE, AT = os.environ["EPISODE_URL"], float(os.environ.get("INSERT_AT", "312.5"))
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
def finish(job):
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
return call(job["result_url"])["result"]
read = finish(call("https://api.sume.com/v1/tts-router/generate", {
"model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
"transcript": "This episode is sponsored by Northwind Coffee. Use code SHOW for ten percent off.",
"output_format": {"container": "wav"}}, "sponsor-read-v1"))
voice = next(a["url"] for a in read["artifacts"] if a["type"] == "audio")
parts = [{"url": EPISODE, "duration": AT}, {"url": voice}, {"url": EPISODE, "source_in": AT}]
try:
job = call("https://api.sume.com/v1/timeline-1.0/audio", {"operation": "concat", "parts": parts}, "sponsor-join-v1")
except urllib.error.HTTPError as e:
print(json.load(e)["error"]["code"]); raise
out = finish(job)
print(out.get("audio_url") or out["artifacts"][0]["url"])Rules the call follows
Three rules from the Timeline audio docs shape the code. Parts are { url, source_in?, duration? }, so the same episode file appears twice with different windows. Every part must share a channel layout, or the job fails with audio_parts_channel_mismatch. And every URL must be this workspace's media.sume.com audio, so import the episode first through POST /v1/media-imports. The call also requires an Idempotency-Key, which the script sets.
Layout, format and length
If the join returns a channel-layout error, your episode and the read differ: a stereo show and a mono voice, or the reverse. Make them match before you join. For an episode that comes from a video, audio detach can output mono; for a stereo episode, check the channel layout of the read's file before you join, and test with a short read first. Audio detach takes channels: "source" (the default) or "mono".
Keep WAV until the very end. The default output is pcm_s16le and sample-exact; MP3 adds priming padding at every edge, so only convert at the last step, with output: { format: "mp3" } on the final join if you want a smaller file. Output is capped at 1,800 seconds, so a long show needs the insert done on the part that contains the break.
| Step | Price | Note |
|---|---|---|
| TTS, 330 characters | $0.0157, billed as $0.02 | $0.0475 per 1,000 characters, rounded up to the cent |
| Timeline audio concat | $0.01 | Flat per job, no provider inference |
| Total | About $0.03 | Per insert, per episode |
Disclosure and scale
If the read is synthetic, say so. Apple's podcast guidelines on disclosing AI voices are summarized in this post. For a host-read ad that you recorded, skip TTS and send your own file as the second part.
To place many reads in many episodes, give every join its own idempotency key, derived from the episode id, the read id and the insert second. A rerun after a crash then returns the finished join instead of paying for a new one. The returned segments[] list the concat offsets, which is useful if you need to publish chapter marks that include the ad.
Pick the insert second from the transcript, not by ear. Run the episode through speech-to-text with sentence segments, find the sentence boundary closest to your intended break, and use that segment's start as INSERT_AT. A cut at a sentence start avoids chopping a word. The same segments give you the timestamp for the chapter marker, so one transcription serves the insert and the show notes.
Sources
Related posts
More in Use cases
- Instagram Live ad clip: burn an offer line with caption cues
Add a timed offer line to a Live replay clip with POST /v1/video-captions and cues, no speech needed. Sume bills $0.20 per clip up to 60 seconds.
- Instagram Live Ads promo: make a 9:16 teaser clip with Sume
Instagram Live Ads started a general rollout on 2026-09-29. Make the 9:16 teaser that sends people to the stream with POST /v1/videos, 3 to 10 seconds.
- Kajabi lesson video from a 16:9 Sume avatar clip: what fits
Kajabi takes MP4 up to 4 GB and recommends 16:9. A 16:9 Sume avatar clip of up to 60 seconds is a lesson intro or recap. The numbers, and where it stops.
- LinkedIn video ad: cut a 3-second bumper from a longer video
LinkedIn's Videos API takes MP4 from 3 seconds to 30 minutes. Cut a 3-second bumper from a finished video with POST /v1/video-trim, $0.02 a job.
Written by Sume