Insert a sponsor read into a podcast: TTS plus one audio concat

Generate the ad read with Sume TTS, then join it into the episode at the second you choose with one Timeline audio concat: sample-exact, $0.01 per join.

5 min readSume
All posts

To insert a sponsor read into an existing podcast episode, generate the read with Sume TTS as a WAV, then call POST /v1/timeline-1.0/audio with operation: "concat" and three parts: the episode up to the insert point, the read, and the episode from the insert point on. The join is sample-domain with no silence at the seams, and costs a flat $0.01. The script below does the whole thing. It needs the episode on media.sume.com, because Timeline audio refuses off-host URLs.

Voice work for ads is getting cheaper to start: the October tracker lists Microsoft's MAI-Voice-2.1 Flash at $15 per 1M characters (read 2026-10-06). A 330-character read is half a cent there. On Sume it is $0.0157 at $0.0475 per 1,000 (pricing), which the Sume price book rounds up to whole cents, so $0.02 on the invoice. The ad read is never the expensive part. Getting it into the episode cleanly is.

Generate and join

import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]

import urllib.error
EPISODE, AT = os.environ["EPISODE_URL"], float(os.environ.get("INSERT_AT", "312.5"))

def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

def finish(job):
    while not job["terminal"]:
        time.sleep(job.get("next_poll_after_seconds") or 2)
        job = call(job["status_url"])
    return call(job["result_url"])["result"]

read = finish(call("https://api.sume.com/v1/tts-router/generate", {
    "model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
    "transcript": "This episode is sponsored by Northwind Coffee. Use code SHOW for ten percent off.",
    "output_format": {"container": "wav"}}, "sponsor-read-v1"))
voice = next(a["url"] for a in read["artifacts"] if a["type"] == "audio")
parts = [{"url": EPISODE, "duration": AT}, {"url": voice}, {"url": EPISODE, "source_in": AT}]
try:
    job = call("https://api.sume.com/v1/timeline-1.0/audio", {"operation": "concat", "parts": parts}, "sponsor-join-v1")
except urllib.error.HTTPError as e:
    print(json.load(e)["error"]["code"]); raise
out = finish(job)
print(out.get("audio_url") or out["artifacts"][0]["url"])

Rules the call follows

Three rules from the Timeline audio docs shape the code. Parts are { url, source_in?, duration? }, so the same episode file appears twice with different windows. Every part must share a channel layout, or the job fails with audio_parts_channel_mismatch. And every URL must be this workspace's media.sume.com audio, so import the episode first through POST /v1/media-imports. The call also requires an Idempotency-Key, which the script sets.

Layout, format and length

If the join returns a channel-layout error, your episode and the read differ: a stereo show and a mono voice, or the reverse. Make them match before you join. For an episode that comes from a video, audio detach can output mono; for a stereo episode, check the channel layout of the read's file before you join, and test with a short read first. Audio detach takes channels: "source" (the default) or "mono".

Keep WAV until the very end. The default output is pcm_s16le and sample-exact; MP3 adds priming padding at every edge, so only convert at the last step, with output: { format: "mp3" } on the final join if you want a smaller file. Output is capped at 1,800 seconds, so a long show needs the insert done on the part that contains the break.

Cost of one sponsor insert (read 2026-10-06)
StepPriceNote
TTS, 330 characters$0.0157, billed as $0.02$0.0475 per 1,000 characters, rounded up to the cent
Timeline audio concat$0.01Flat per job, no provider inference
TotalAbout $0.03Per insert, per episode

Disclosure and scale

If the read is synthetic, say so. Apple's podcast guidelines on disclosing AI voices are summarized in this post. For a host-read ad that you recorded, skip TTS and send your own file as the second part.

To place many reads in many episodes, give every join its own idempotency key, derived from the episode id, the read id and the insert second. A rerun after a crash then returns the finished join instead of paying for a new one. The returned segments[] list the concat offsets, which is useful if you need to publish chapter marks that include the ad.

Pick the insert second from the transcript, not by ear. Run the episode through speech-to-text with sentence segments, find the sentence boundary closest to your intended break, and use that segment's start as INSERT_AT. A cut at a sentence start avoids chopping a word. The same segments give you the timestamp for the chapter marker, so one transcription serves the insert and the show notes.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume