Audio podcast to video with AI: cover art plus the audio

Turn podcast audio into a video by holding one cover image for the whole episode on Timeline 1.0, then cut short clips with an audio split.

5 min readSume
All posts

You can turn podcast audio into a video by holding one still image, usually the cover art, on screen for the whole episode with the audio as the soundtrack. On Sume that is one Timeline 1.0 render: the audio is the spine, and a single still fills the picture.

The audio and the image must already be hosted on Sume's media.sume.com, such as the output of an earlier Sume job. Details are from Timeline 1.0 and Timeline audio, read 2026-09-29. To make the audio itself, see AI podcast generator from text.

How do I put cover art over an audio file?

Send POST /v1/timeline-1.0/render with the audio as audio.url, its length as audio.duration_seconds, and one video[] slot whose source_url is the still and whose duration matches the audio. Stills are static holds, and the first slot starts at 0. Cover art is usually square, so set output to a square frame; the default is 1080×1920 and fit defaults to cover, which crops.

An image made with POST /v1/images comes back as a Sume-hosted media.sume.com URL, so cover art you generate there can go straight into the slot. For the size Apple asks for on cover art, see podcast cover art size.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: podcast-video-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/episode.mp3",
      "duration_seconds": 300
    },
    "output": { "width": 1080, "height": 1080, "fps": 30 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/cover.png",
        "start": 0, "duration": 300 }
    ]
  }'

What are the limits for a podcast video?

The render is a job: poll GET /v1/jobs/:id/status, then read GET /v1/jobs/:id/result for the video_url.

From Timeline 1.0 and Timeline audio, read 2026-09-29.
FieldRule
audio.duration_secondsRequired, 1 to 1800 seconds
video[] slots1 to 200; video[0].start must be 0
video[].durationAt least 0.2 s; may trail the spine by at most 0.5 s
output.width / heightEven integers, 256 to 2160
Audio split ranges[]1 to 20 per job
Audio split outputwav by default, or mp3

How do I cut short clips from an episode?

Run POST /v1/timeline-1.0/audio with operation: "split", the episode's Sume-hosted url, and ranges of { start, end } seconds. Each range returns its own media.sume.com audio_url, with segments[] timings. Then render one video per clip with the same still, using the clip as audio.url.

To find where the good moments are, speech to text gives the words; STT 1.0 takes at most 10 minutes per request, so a long episode goes in parts.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: podcast-split-001" \
  -d '{
    "operation": "split",
    "url": "https://media.sume.com/artifacts/artf_demo/episode.wav",
    "ranges": [{ "start": 62.5, "end": 95 }, { "start": 410, "end": 452.3 }]
  }'

Can I add captions or a waveform?

Timeline 1.0's program table has no waveform or audio-visualizer field, so a Sume podcast video is a still plus sound. Captions are a separate job: POST /v1/video-captions burns them onto a finished video URL, and speech-to-captions needs audible speech. In current code the caption job refuses a source over 60 seconds, so captions suit short clips, not a full episode.

How much does a podcast video cost?

A Timeline render is billed at $0.10 per output minute on the API pricing page, with no model inference. A 300-second episode therefore bills five output minutes. The audio split is a separate flat per-job fee; read it on the pricing page or from GET /v1/catalog.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume