AI play-by-play voiceover for a sports highlight reel: how to time it
Write one line per play, ask TTS for sentence timings, and place each clip on its line: a 60-second reel costs about 14 cents on Sume. Not live commentary.

To voice a sports highlight reel with AI, write one sentence per play, send the whole script as one TTS job with word timestamps and sentence segmentation turned on, and use the returned sentence timings as the start times of the clips in a Timeline render. For six plays of about 125 characters, the voice is $0.04 and the 60-second render is $0.10, so about $0.14 for the reel on Sume at the rates in the catalog on 2026-10-07.
This is post-production for footage you own or have the rights to. Sume's TTS is asynchronous and has no streaming output, so it cannot call a live match; it can voice a recap a few minutes after the final whistle.
Write the script as one line per play
Cartesia's pricing page puts one minute of Sonic speech at 750 to 800 credits, one per character, and Sume TTS 1.0 meters one credit per character. A 60-second reel at about 750 characters a minute holds roughly 750 characters of narration if the voice talks the whole time. Six lines of about 125 characters is 750 characters, so the voice fills the reel; if you want breathing room for action sound, write five lines or raise generation_config.speed toward 1.15. Time one test line before you commit to the count.
Keep each line to one sentence and avoid abbreviations the engine might spell out. Put the player name in the sentence where the clip shows him or her, since the timing for each sentence is what places the footage. At $0.0475 per 1,000 characters, 750 characters is $0.0356, rounded up to $0.04.
Ask for sentence timings in wav
Set timestamps.words to true and segmentation.mode to sentence. The result then carries gapless segments[] where each segment ends exactly where the next begins; the default cut is 70 ms after the last word, and the next segment absorbs the pause. With a wav or raw container each segment also gets its own sample-exact audio_url; with mp3 you get the timings only.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: highlights-001" \
-d @- <<JSON
{
"transcript": "Eleven minutes in, the striker cuts inside and curls it into the far corner. Two minutes later the keeper answers with a full-stretch save.",
"voice": { "id": "$VOICE_ID" },
"language": "en",
"generation_config": { "speed": 1.1, "emotion": "excited, fast, broadcast" },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence", "emit_audio": true },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"mode": "async"
}
JSONPlace the clips on the spine
Timeline 1.0 takes the voice file as audio and an ordered video[] of up to 200 slots. The first slot must start at 0 and later starts must increase, and coverage may stop at most 0.5 seconds before the end of the audio. Give each clip the start time of its sentence, and set its duration to the sentence length. Short clips are padded or looped and reported as soft warnings, so check the render's warnings[].
Import the clips first with a media import; every URL in a render must be a media.sume.com file in your workspace. The default output is a 1080 by 1920 vertical MP4 and the render is $0.10 per started output minute. Add a soundtrack with duck_db if you want crowd noise or a stadium-style bed under the voice; generate the bed with a Music Router prompt for $0.125.
Bill of materials
- A re-voice of one changed line costs a full job for the whole script, unless you split the script into jobs of a few plays each.
- Per-sentence wav slices let you swap one line's audio without regenerating the others, as long as you keep the same sentence length or re-time the clip.
| Step | Unit | Price | This reel |
|---|---|---|---|
| TTS 1.0 narration, 750 characters | per 1,000 characters | $0.0475 | $0.04 |
| Timeline 1.0 render, 60 seconds | per started output minute | $0.10 | $0.10 |
| Optional crowd bed from Music Router | per generation | $0.125 | $0.125 |
| Total without the bed | $0.14 | ||
| Total with the bed | $0.27 |
Limits worth knowing before the weekend
Voices are chosen by voice id or by an avatar whose voice is ready; there is no picker that says 'commentator'. Audition two or three voices on the first sentence, then fix the one you like and reuse its id for every reel so the channel keeps one sound. Names of players and clubs are the usual mispronunciation risk; a pronunciation dictionary id can be attached to a request, and a 5-second test of the roster is cheaper than a retake.
MAI-Voice-2.1's model page puts its Flash tier at about 45 ms latency and the standard tier at about 550 ms; neither matters for a recap, where the file is ready minutes before the clip edit is.
Sources
Related posts
More in Use cases
- AI spokesperson release notes, 90 seconds: two jobs, cost by tier
A 90-second spokesperson video is two Sume avatar jobs because one job tops out at 60 seconds. Cost at standard, plus and max, and how to split the script.
- How do I make a spoken wake-up alarm with an AI voice and music?
A wake-up alarm is a short voice greeting joined to a music file: 2 cents of TTS and a 1-cent concat per day, with one $0.125 music track reused all week.
- Animate a painting with AI video: first frame and a slow camera prompt
Pin a painting as the first frame with frame_images on Sume, ask for a slow camera move, and pick the model by length and price. A 6-second price table.
- Animate a pencil sketch or storyboard panel as the first frame
Use a sketch or storyboard panel as first_frame on Sume, and the next panel as last_frame on models that take an end frame. Prompt, models and windows.
Written by Sume