Podcast audiogram API: audio plus a cover still in one render
Make a podcast audiogram video from an audio clip and a cover image with Timeline 1.0: spine, still slot, square output and the cost per minute.

A podcast audiogram is a video whose picture is a still cover image and whose sound is a clip of the episode. In Sume that is one Timeline 1.0 render: the audio goes on the spine, a single image fills video[] for the same length, and you set a square or vertical output. A 60-second clip costs $0.10, because Timeline bills $0.10 per output minute, rounded up.
This is the plain version. It does not draw an animated waveform, and we say so up front so you do not plan around one.
The request
Upload the episode audio and the cover image first (POST /v1/assets/upload-url, PUT the bytes, then POST /v1/assets/{id}/complete; the ready asset's url is the media.sume.com URL Timeline accepts); Timeline only accepts your workspace's media.sume.com URLs. audio.duration_seconds is required, and it is the output length. To cut a clip from the middle of a long episode, point audio.url at the full file and set audio.source_in to the in-point; the output is still duration_seconds long. Still images are static holds, and the slot must cover the spine to within half a second.
Plan it first with the unbilled POST /v1/timeline-1.0/plan, then render with an Idempotency-Key.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audiogram-ep42-clip1" \
-d '{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/episode.wav",
"source_in": 754.5,
"duration_seconds": 60
},
"video": [{
"source_url": "https://media.sume.com/artifacts/artf_demo/cover.png",
"start": 0,
"duration": 60,
"fit": "cover"
}],
"output": { "width": 1080, "height": 1080, "fade_out_seconds": 1 }
}'What it costs and allows
| Item | Value |
|---|---|
| Rate | $0.10 per output minute, rounded up |
| 60-second clip | $0.10 |
| 10-minute full episode | $1.00 |
| Longest spine | 1800 seconds (30 minutes) |
| Default output | 1080 x 1920; set width and height for square |
| Still image slot | Static hold; motion is accepted and ignored |
Add captions as a second step
A silent cover is a weak post, so burn captions on the rendered MP4 with POST /v1/video-captions. Speech-to-captions is $0.20 per accepted job and needs audible speech, which an audiogram has. Use style: "slam" or black-outline and pass script_text if you have a transcript, so names are spelled your way. The caption step re-encodes the video, so do it last.
Limits
There is no waveform, no progress bar and no speaker-turn graphics in this endpoint. If you want a moving picture, generate a few seconds of looping visual with a video model and put it in video[] instead of the still; the audio spine does not change. A source longer than 30 minutes needs source_in clips, not one render. Check each platform's own file limits before upload; those are covered in the platform posts, not here.
Sources
Related posts
More in Use cases
- Turn a product changelog into a 30-second update video
Convert release notes into a short update video: pick three changes, write a 70-word script, voice it with TTS, add real screenshots and render with Timeline.
- Launch teaser from product shots: Gemini Omni Flash reference images
Feed up to 10 product references to gemini-omni-flash-1.1 on Sume, address them as IMAGE_REF_0 in the prompt, and get a 3-10 second teaser with native audio.
- Product photo on a white background: one image edit call on Sume
Turn a messy product photo into a clean white-background shot with one POST /v1/images edit. Which model, which fields, and what to check before you publish.
- Re-roll one shot in a stitched film without redoing the rest
Regenerate a single bad shot, swap its source_url in the Timeline 1.0 video array, and re-render the cut for $0.10 per output minute. The other shots stay.
Written by Sume