Podcast audiogram API: audio plus a cover still in one render

Make a podcast audiogram video from an audio clip and a cover image with Timeline 1.0: spine, still slot, square output and the cost per minute.

5 min readSume
All posts

A podcast audiogram is a video whose picture is a still cover image and whose sound is a clip of the episode. In Sume that is one Timeline 1.0 render: the audio goes on the spine, a single image fills video[] for the same length, and you set a square or vertical output. A 60-second clip costs $0.10, because Timeline bills $0.10 per output minute, rounded up.

This is the plain version. It does not draw an animated waveform, and we say so up front so you do not plan around one.

The request

Upload the episode audio and the cover image first (POST /v1/assets/upload-url, PUT the bytes, then POST /v1/assets/{id}/complete; the ready asset's url is the media.sume.com URL Timeline accepts); Timeline only accepts your workspace's media.sume.com URLs. audio.duration_seconds is required, and it is the output length. To cut a clip from the middle of a long episode, point audio.url at the full file and set audio.source_in to the in-point; the output is still duration_seconds long. Still images are static holds, and the slot must cover the spine to within half a second.

Plan it first with the unbilled POST /v1/timeline-1.0/plan, then render with an Idempotency-Key.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audiogram-ep42-clip1" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/episode.wav",
      "source_in": 754.5,
      "duration_seconds": 60
    },
    "video": [{
      "source_url": "https://media.sume.com/artifacts/artf_demo/cover.png",
      "start": 0,
      "duration": 60,
      "fit": "cover"
    }],
    "output": { "width": 1080, "height": 1080, "fade_out_seconds": 1 }
  }'

What it costs and allows

Timeline 1.0 limits and rate from docs.sume.com/models/timeline, read 2026-10-01.
ItemValue
Rate$0.10 per output minute, rounded up
60-second clip$0.10
10-minute full episode$1.00
Longest spine1800 seconds (30 minutes)
Default output1080 x 1920; set width and height for square
Still image slotStatic hold; motion is accepted and ignored

Add captions as a second step

A silent cover is a weak post, so burn captions on the rendered MP4 with POST /v1/video-captions. Speech-to-captions is $0.20 per accepted job and needs audible speech, which an audiogram has. Use style: "slam" or black-outline and pass script_text if you have a transcript, so names are spelled your way. The caption step re-encodes the video, so do it last.

Limits

There is no waveform, no progress bar and no speaker-turn graphics in this endpoint. If you want a moving picture, generate a few seconds of looping visual with a video model and put it in video[] instead of the still; the audio spine does not change. A source longer than 30 minutes needs source_in clips, not one render. Check each platform's own file limits before upload; those are covered in the platform posts, not here.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume