Add a spoken call-to-action outro to a finished video, with cost
Generate a 3 to 5 second CTA line with Sume TTS and append it to a finished video with Timeline, keeping the video's own sound. Steps, code, and cost.

To add a spoken call-to-action to the end of a finished video and keep its sound, make three things: detach the video's own audio, generate the CTA line with TTS, and render one Timeline job whose audio is those two files joined in order, with the clip on screen and then an end card. Timeline 1.0 is driven by its audio spine, so the original sound has to be declared again, as the first part of that spine. For a clip under a minute the total is about $0.113: $0.01 detach, $0.10 render, and under a cent for a 60-character line.
The endpoints and rules come from Timeline 1.0, Timeline audio, Audio detach, the Sume API reference and API pricing, read on 2026-10-03. A voiceover mixed across a whole video is Add a voiceover to a video; a branded intro and outro made of video files is Add an intro and outro to every clip; this page is one spoken line after the end.
How long should the CTA line be?
Write 40 to 60 characters, one sentence, and ask for timestamps.words so the result reports duration_seconds. Sume's docs give no characters-per-second rule, so measure the take and aim for 3 to 5 seconds. Generate it as wav at 44100 Hz so the join is sample exact. At $0.0475 per 1,000 characters, 60 characters cost $0.00285.
How do I attach it so the video keeps its sound?
Both files must already be on media.sume.com. First, POST /v1/audio-detach on the clip, which gives a wav and its duration_seconds. Then render with audio.parts set to the detached wav and the CTA wav. In current code the TTS wav is one channel, so detach with channels: "mono" to match, or the join can fail with audio_parts_channel_mismatch. The spine joins gaplessly, with no silence at the seam. Declare audio.duration_seconds as the clip length plus the CTA length. The clip is slot one; slot two is a still end card on screen for the CTA, starting at the clip's length. Match output to the clip's size, since the default is 1080 by 1920.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: outro-cta-001" \
-d '{
"audio": {
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/clip-audio.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/cta.wav" }
],
"duration_seconds": 45.6
},
"output": { "width": 1080, "height": 1920 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 0, "duration": 42 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/end-card.png", "start": 42, "duration": 3.6 }
]
}'What does the outro cost?
The render is billed at $0.10 per output minute, rounded up, so crossing a minute line costs another $0.10.
- A 58-second clip plus a 4-second line is 62 seconds, which rounds up to two minutes and $0.20 for the render.
- Run
POST /v1/timeline-1.0/planfirst. It is unbilled and returnsbillable_minutes. - A mono spine under stereo sources can raise an
audio_spine_low_fidelitywarning; readwarnings[].
| Step | Endpoint | Price |
|---|---|---|
| CTA line, 60 characters | POST /v1/tts-1.0/generate | $0.00285 |
| Detach the clip's audio | POST /v1/audio-detach | $0.01 |
| Render, 45.6 seconds | POST /v1/timeline-1.0/render | $0.10 |
| Total | Three jobs | About $0.113 |
What if the picture and the sound drift apart?
Take the clip's length from the detach result's duration_seconds, use it for slot one's duration and slot two's start, and add the CTA's duration_seconds for the total. Video coverage may trail the spine by at most 0.5 seconds, so make the slot durations add up to audio.duration_seconds. Read warnings[] on the result for padded or looped sources.
Sources
Related posts
More in Media tools
- Add captions to a recorded AI avatar call: what Sume needs
Tavus can record a live avatar call to your bucket. Sume's video captions need a public HTTPS URL and priced for clips up to 60 seconds. Steps and limits.
- Add subtitles to a video in DaVinci Resolve, or with a caption API
Resolve 21 transcribes and translates speech to text in the app. Sume video-captions burns styled captions into a public MP4 URL for $0.20 up to 60 seconds.
- AI music from an image: Lyria 3.5 image_url mood still on Sume
Sume's Music Router takes an optional public HTTPS image_url as a mood still next to the prompt. How to send it, what it costs and what to expect back.
- AI music generator for covers: what Sume cannot do, and a swap
Sume cannot cover an existing song: no audio input and no song-to-song path. Licensed cover platforms are still in development. Here is an original-track swap.
Written by Sume