Create motivational videos with AI: voice, shots, and music

Create motivational videos with AI: a slow spoken quote over cinematic shots, a music bed that builds and ducks under the voice, and big captions.

5 min readSume
All posts

To create a motivational video with AI, voice one short line, a quote or a lesson, a little slower than normal speech, and lay it over two or three cinematic shots with a music bed that builds and big captions. Each part is its own AI step: text to speech for the voice, a video model for the shots, a music model for the bed, and one render that lowers the music under the voice.

The settings below come from the Music Router, Music 1.0, Video generation, Timeline 1.0, and Video captions docs and the TTS schema in the Sume API reference, read on 2026-09-28. Automate faceless short-form videos covers the general pipeline; this post is the motivational version.

How do I make the voice sound motivational?

Pace does most of the work. Send the line to TTS 1.0 with generation_config.speed below 1 so every word lands, and try the optional emotion guide. The schema lists no accepted values for emotion, so test a few words, such as calm or determined, and listen. Use a quote only if you can attribute it correctly and have the right to use it.

Delivery and mix settings from the Sume API reference and Timeline 1.0, read 2026-09-28.
SettingRangeUse in a motivational video
TTS generation_config.speed0.6–1.5Below 1 for a slower, weighted read
TTS generation_config.emotionFree text, up to 64 charactersAn optional guide; test it
Timeline soundtrack.duck_db0–20How far the music dips while the voice speaks
Timeline soundtrack.fade_out_seconds0–10Lets the music ring out at the end

How do I make the music build?

Describe the arc in a music brief. Sume's music docs suggest seven axes: a precise emotion, a genre, the tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and an era. Close with "Instrumental, no vocals.", adding "no spoken word" under narration. The docs call these creative directions, not guaranteed settings, so listen to the result.

Send the brief to POST /v1/music-router/generate. There is no duration field: duration and duration_seconds are rejected, so ask for the length in the prompt. The result's audio artifact is Sume-hosted, ready for the render. Music is $0.125 per audio. Generate music for a video with AI shows how to fit a track to a cut.

Determined, slowly rising cinematic score, 70 BPM, D minor.
Low strings and a soft piano ostinato, distant taiko.
Quiet for the first 10 seconds; full strings and drums arrive at 0:12.
A 30-second track. Instrumental, no vocals, no spoken word.

How do I put the voice, shots, and music together?

Generate two or three shots with POST /v1/videos at aspect_ratio: "9:16": a lone runner at dawn, a mountain ridge, city lights. Then one Timeline 1.0 render puts them over the voice as the audio spine, with the music as soundtrack; in current code those two are the only sound in the render, and the audio of each clip is dropped. Omit soundtrack.gain_db and the bed sits at −16 dB; duck_db dips it further while the voice speaks, and it comes back up in the pauses. What is audio ducking? explains the effect. The render's default output is 1080×1920, and every URL must be a Sume-hosted file from an earlier Sume job.

How do I add big, bold captions?

Send the rendered video to POST /v1/video-captions with style: "slam", which in current code shows one uppercase word at a time; Hormozi-style captions describes that look. Name it: an omitted style also resolves to slam for Latin-script wording such as English, but the docs say a defaulted render drops the gold accent on the spoken word. To skip speech-to-text and burn exactly your quote, send the TTS word timings as words. Each entry needs text, start, and end, and in current code end must be after start, so rename the TTS result's word field to text. In current code a caption job refuses a video over 60 seconds, so keep each video under a minute.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume