Spotify verified podcast trailer from your own audio, no clone

Spotify says it removes shows that clone a host's voice without permission. Cut a trailer from your own episode audio with Sume Timeline, no synthetic voice.

4 min readSume
All posts

A podcast trailer built from your own recording needs no synthetic voice. Take the real episode audio, pick the best 30 to 60 seconds, set it as the audio.url of a Sume Timeline render with source_in at the moment you want, put your cover art or a clip on screen, and add captions. The result is a trailer in your host's actual voice, which matters this year because Spotify says it will remove podcasts that impersonate a host's likeness, including with AI voice cloning.

What did Spotify say about verification and cloning?

Spotify's newsroom post describes a "Verified by Spotify" badge for podcasts: a light green checkmark that marks a show as the official presence of a creator, publisher or brand. The criteria it lists are consistent audience engagement over time, good standing with platform policies, and verified audience authenticity. Spotify says the rollout began in May 2026 on select shows and continues over coming months.

The post also says Spotify will remove shows and content that impersonate another creator or host's likeness without permission, whether through AI voice cloning or any other method. Nothing on the page says a trailer cut from your own audio is affected. The point for a show owner is simple: ship your trailer in your own recorded voice.

How does the audio spine take a slice of an episode?

Timeline 1.0 has audio.source_in, the in-point into a single url spine, with the output length still set by duration_seconds (1 to 1800). So a 45-second trailer from minute 12 is source_in: 720 and duration_seconds: 45. The audio must already be a media.sume.com file. If the episode only exists as a video, pull the sound out with audio detach, a $0.01 job that returns a durable audio file.

Trailer fields from the Timeline 1.0 docs and prices from the Sume docs, read 2026-10-03.
NeedField or surfaceValue
Slice of the episodeaudio.source_in with audio.urlSeconds into the file
Trailer lengthaudio.duration_seconds1 to 1800
Cover art on screenvideo[] still slotStatic hold, duration of at least 0.2 s
Audio from a video fileAudio detach$0.01 per job
RenderTimeline 1.0$0.10 per output minute

How do I render it?

Use the cover as a still slot for the full length, or a few host clips from the recording if you have them.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: trailer-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_ep/episode-12.wav",
      "source_in": 720,
      "duration_seconds": 45
    },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_cover/cover.jpg", "start": 0, "duration": 45 }
    ]
  }'

How do I caption the trailer?

Run video captions on the render's video_url. Because there is real speech, you can leave out cues and let speech-to-text produce the words, or pass script_text if you want exact wording. The slam style is the default for Latin text; set language when the show is not English. Read the caption job's result for the final file.

What stays outside this workflow?

Sume does not apply a Verified badge or file anything with Spotify. Verification is a Spotify decision on audience and policy criteria, and the trailer does not change it. The only thing the trailer touches is how you represent your host: with a real recording, you keep the voice honest. If you do use a synthetic voice anywhere in the show, label it as such and keep the host's permission on file.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume