How do I add a voiceover and music to a timelapse video?
Narrate a 45-second timelapse with one TTS job, one music track and one Timeline render: 3 cents, $0.125 and $0.10, about 33 cents on Sume.

To add a voiceover and music to a timelapse, write a script that fits the clip's length, generate it with one TTS job, generate a music bed with Music Router, and lay both over the clip in a Timeline 1.0 render. For a 45-second timelapse the voice is about 560 characters, $0.03 on Sume; the bed is $0.125; the render is $0.10 for the first minute. Total: $0.255 at the catalog rates of 2026-10-07.
A timelapse is a good fit because its length is fixed and its picture is silent. The voice becomes the audio spine of the render, the clip is the only video slot, and the music sits underneath with a duck so the speech stays clear.
Fit the script to the clip
Cartesia's pricing page gives one minute of Sonic speech as 750 to 800 credits at one credit per character, and Sume TTS 1.0 meters one credit per character. For a 45-second clip, 560 characters at 750 a minute is 45 seconds, which leaves no gap at all. Aim for 450 to 500 characters and let the picture breathe at the start and the end.
If the first take is long, change generation_config.speed instead of cutting words: the range is 0.6 to 1.5 and the price depends on characters, not seconds, so a faster read costs the same as a slower one. Measure the finished file's duration before you set the render's audio.duration_seconds.
The voice and the bed
Generate the voice as wav so the render does not stack encoder padding on top of it. For the bed, ask Music Router for a track whose mood follows the subject: a building site wants steady and forward, a sunset wants patient. Put the length in the prompt, since a duration field is rejected, and end with 'Instrumental, no vocals' so the bed leaves room for words. Router prompts run 1 to 5,000 characters and every router model charges the same $0.125.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timelapse-bed-001" \
-d '{
"model": "sume/music-auto",
"prompt": "A 50-second patient, forward-moving ambient piece, 92 BPM, D major. Warm pads, soft pulsing bass, a single piano motif returning every eight bars. Instrumental, no vocals."
}'The render
Timeline 1.0 wants Sume-hosted URLs, so import your timelapse first with a media import. Then send the voice as audio.url, the timelapse as the single video[] slot, and the bed as the soundtrack with duck_db between 0 and 20 and a fade_out_seconds up to 10. Set the slot duration to the voice length so the picture ends with the words. Check the result's warnings[] before you publish.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timelapse-render-001" \
-d '{
"audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 45 },
"video": [{ "source_url": "https://media.sume.com/artifacts/artf_demo/timelapse.mp4", "start": 0, "duration": 45, "fit": "cover" }],
"soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": -8, "duck_db": 10, "fade_out_seconds": 3 },
"output": { "width": 1080, "height": 1920 }
}'Pacing the script against the picture
Timelapses move in beats: a sky darkens, a crane lifts a beam, a floor appears. Write one sentence per beat and keep each under 120 characters, so the voice names the change a second before the viewer sees it. A cold open that states the total span, such as 'Eleven weeks, one corner lot', gives the clip a frame before the first beat.
Avoid filler adjectives and numbers that the picture can show. Spoken digits cost characters and listeners lose them. Spell out units the way you want them said, and use a pronunciation dictionary if a place name or brand is read wrongly; the retake is a new job at the same rate.
Finally, play the clip once without the bed. If the voice still carries the story, the bed only needs to add warmth, and a duck of 8 to 12 dB will be plenty.
Cost and what to check
- A timelapse often has no sound of its own, so the render's audio is only your voice and bed; no original audio to preserve.
- Match the aspect: the default output is 1080 by 1920, so set
output.widthandheightfor a landscape clip. - Retake the voice alone if a word is wrong: that is another $0.03 and a re-render at $0.10.
- Do not let the voice describe what the picture already shows. Say why it matters.
| Step | Unit | Price | This clip |
|---|---|---|---|
| TTS 1.0, 560 characters | per 1,000 characters | $0.0475 | $0.03 |
| Music Router bed | per generation | $0.125 | $0.125 |
| Timeline 1.0 render, 45 s | per started output minute | $0.10 | $0.10 |
| Total | $0.255 |
Sources
Related posts
More in Use cases
- Add sunglasses or a hat to a person in an existing video (Omni edit)
Use Gemini Omni Flash 1.1 video_url edit on Sume to add an accessory to a person already in a clip. Prompt wording, a frame check list, and 4-second prices.
- AI play-by-play voiceover for a sports highlight reel: how to time it
Write one line per play, ask TTS for sentence timings, and place each clip on its line: a 60-second reel costs about 14 cents on Sume. Not live commentary.
- AI spokesperson release notes, 90 seconds: two jobs, cost by tier
A 90-second spokesperson video is two Sume avatar jobs because one job tops out at 60 seconds. Cost at standard, plus and max, and how to split the script.
- How do I make a spoken wake-up alarm with an AI voice and music?
A wake-up alarm is a short voice greeting joined to a music file: 2 cents of TTS and a 1-cent concat per day, with one $0.125 music track reused all week.
Written by Sume