AI play-by-play voiceover for a sports highlight reel: how to time it

Write one line per play, ask TTS for sentence timings, and place each clip on its line: a 60-second reel costs about 14 cents on Sume. Not live commentary.

4 min readSume
All posts

To voice a sports highlight reel with AI, write one sentence per play, send the whole script as one TTS job with word timestamps and sentence segmentation turned on, and use the returned sentence timings as the start times of the clips in a Timeline render. For six plays of about 125 characters, the voice is $0.04 and the 60-second render is $0.10, so about $0.14 for the reel on Sume at the rates in the catalog on 2026-10-07.

This is post-production for footage you own or have the rights to. Sume's TTS is asynchronous and has no streaming output, so it cannot call a live match; it can voice a recap a few minutes after the final whistle.

Write the script as one line per play

Cartesia's pricing page puts one minute of Sonic speech at 750 to 800 credits, one per character, and Sume TTS 1.0 meters one credit per character. A 60-second reel at about 750 characters a minute holds roughly 750 characters of narration if the voice talks the whole time. Six lines of about 125 characters is 750 characters, so the voice fills the reel; if you want breathing room for action sound, write five lines or raise generation_config.speed toward 1.15. Time one test line before you commit to the count.

Keep each line to one sentence and avoid abbreviations the engine might spell out. Put the player name in the sentence where the clip shows him or her, since the timing for each sentence is what places the footage. At $0.0475 per 1,000 characters, 750 characters is $0.0356, rounded up to $0.04.

Ask for sentence timings in wav

Set timestamps.words to true and segmentation.mode to sentence. The result then carries gapless segments[] where each segment ends exactly where the next begins; the default cut is 70 ms after the last word, and the next segment absorbs the pause. With a wav or raw container each segment also gets its own sample-exact audio_url; with mp3 you get the timings only.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: highlights-001" \
  -d @- <<JSON
{
  "transcript": "Eleven minutes in, the striker cuts inside and curls it into the far corner. Two minutes later the keeper answers with a full-stretch save.",
  "voice": { "id": "$VOICE_ID" },
  "language": "en",
  "generation_config": { "speed": 1.1, "emotion": "excited, fast, broadcast" },
  "timestamps": { "words": true },
  "segmentation": { "mode": "sentence", "emit_audio": true },
  "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
  "mode": "async"
}
JSON

Place the clips on the spine

Timeline 1.0 takes the voice file as audio and an ordered video[] of up to 200 slots. The first slot must start at 0 and later starts must increase, and coverage may stop at most 0.5 seconds before the end of the audio. Give each clip the start time of its sentence, and set its duration to the sentence length. Short clips are padded or looped and reported as soft warnings, so check the render's warnings[].

Import the clips first with a media import; every URL in a render must be a media.sume.com file in your workspace. The default output is a 1080 by 1920 vertical MP4 and the render is $0.10 per started output minute. Add a soundtrack with duck_db if you want crowd noise or a stadium-style bed under the voice; generate the bed with a Music Router prompt for $0.125.

Bill of materials

  • A re-voice of one changed line costs a full job for the whole script, unless you split the script into jobs of a few plays each.
  • Per-sentence wav slices let you swap one line's audio without regenerating the others, as long as you keep the same sentence length or re-time the clip.
A 60-second, six-play highlight reel on Sume, rates as of 2026-10-07 (TTS 1.0, Timeline 1.0 and Music Router catalog entries); vendor figures read 2026-10-07.
StepUnitPriceThis reel
TTS 1.0 narration, 750 charactersper 1,000 characters$0.0475$0.04
Timeline 1.0 render, 60 secondsper started output minute$0.10$0.10
Optional crowd bed from Music Routerper generation$0.125$0.125
Total without the bed$0.14
Total with the bed$0.27

Limits worth knowing before the weekend

Voices are chosen by voice id or by an avatar whose voice is ready; there is no picker that says 'commentator'. Audition two or three voices on the first sentence, then fix the one you like and reuse its id for every reel so the channel keeps one sound. Names of players and clubs are the usual mispronunciation risk; a pronunciation dictionary id can be attached to a request, and a 5-second test of the roster is cheaper than a retake.

MAI-Voice-2.1's model page puts its Flash tier at about 45 ms latency and the standard tier at about 550 ms; neither matters for a recap, where the file is ready minutes before the clip edit is.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume