Create motivational videos with AI: voice, shots, and music
Create motivational videos with AI: a slow spoken quote over cinematic shots, a music bed that builds and ducks under the voice, and big captions.

To create a motivational video with AI, voice one short line, a quote or a lesson, a little slower than normal speech, and lay it over two or three cinematic shots with a music bed that builds and big captions. Each part is its own AI step: text to speech for the voice, a video model for the shots, a music model for the bed, and one render that lowers the music under the voice.
The settings below come from the Music Router, Music 1.0, Video generation, Timeline 1.0, and Video captions docs and the TTS schema in the Sume API reference, read on 2026-09-28. Automate faceless short-form videos covers the general pipeline; this post is the motivational version.
How do I make the voice sound motivational?
Pace does most of the work. Send the line to TTS 1.0 with generation_config.speed below 1 so every word lands, and try the optional emotion guide. The schema lists no accepted values for emotion, so test a few words, such as calm or determined, and listen. Use a quote only if you can attribute it correctly and have the right to use it.
| Setting | Range | Use in a motivational video |
|---|---|---|
TTS generation_config.speed | 0.6–1.5 | Below 1 for a slower, weighted read |
TTS generation_config.emotion | Free text, up to 64 characters | An optional guide; test it |
Timeline soundtrack.duck_db | 0–20 | How far the music dips while the voice speaks |
Timeline soundtrack.fade_out_seconds | 0–10 | Lets the music ring out at the end |
How do I make the music build?
Describe the arc in a music brief. Sume's music docs suggest seven axes: a precise emotion, a genre, the tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and an era. Close with "Instrumental, no vocals.", adding "no spoken word" under narration. The docs call these creative directions, not guaranteed settings, so listen to the result.
Send the brief to POST /v1/music-router/generate. There is no duration field: duration and duration_seconds are rejected, so ask for the length in the prompt. The result's audio artifact is Sume-hosted, ready for the render. Music is $0.125 per audio. Generate music for a video with AI shows how to fit a track to a cut.
Determined, slowly rising cinematic score, 70 BPM, D minor.
Low strings and a soft piano ostinato, distant taiko.
Quiet for the first 10 seconds; full strings and drums arrive at 0:12.
A 30-second track. Instrumental, no vocals, no spoken word.How do I put the voice, shots, and music together?
Generate two or three shots with POST /v1/videos at aspect_ratio: "9:16": a lone runner at dawn, a mountain ridge, city lights. Then one Timeline 1.0 render puts them over the voice as the audio spine, with the music as soundtrack; in current code those two are the only sound in the render, and the audio of each clip is dropped. Omit soundtrack.gain_db and the bed sits at −16 dB; duck_db dips it further while the voice speaks, and it comes back up in the pauses. What is audio ducking? explains the effect. The render's default output is 1080×1920, and every URL must be a Sume-hosted file from an earlier Sume job.
How do I add big, bold captions?
Send the rendered video to POST /v1/video-captions with style: "slam", which in current code shows one uppercase word at a time; Hormozi-style captions describes that look. Name it: an omitted style also resolves to slam for Latin-script wording such as English, but the docs say a defaulted render drops the gold accent on the spoken word. To skip speech-to-text and burn exactly your quote, send the TTS word timings as words. Each entry needs text, start, and end, and in current code end must be after start, so rename the TTS result's word field to text. In current code a caption job refuses a video over 60 seconds, so keep each video under a minute.
Sources
Related posts
More in Use cases
- Press release video: an announcement read by an AI avatar
A press release video reads the headline, key facts, and a quote in about a minute, captioned from the release text. How to make one with an AI avatar.
- New product launch video with an AI avatar presenter
A new product launch video names the problem, shows what's new, and ends on one call to action. Make it from launch copy with an AI avatar presenter.
- Product URL to video AI: how link-to-video tools work
Product URL to video AI reads a page's title, copy and images, then scripts and renders an ad. How it works, and what to send when a tool can't.
- Product video prompt examples: a template and 6 shots
A product video prompt describes one shot: the action, one camera move, the light, the setting. A template, six examples, and what goes in fields.
Written by Sume