Newsletter to video: sentence-timed TTS drives four slots, 74 cents

Turn a 900-character newsletter intro into a video on Sume: sentence segments set each slot start, four stills, one render, one caption job (read 2026-10-10).

4 min readSume
All posts

A 900-character newsletter intro becomes a captioned four-scene video for about $0.74 on Sume: TTS ($0.0428), four stills ($0.40), one Timeline render ($0.10) and one caption job ($0.20). The sentence timings from the TTS result decide when each still appears, so you do not guess the cut points.

The point of the workflow is that one request returns both audio and timing. Text-to-speech can segment by sentence and return gapless segments, and Timeline takes the audio as its spine, so each sentence group becomes one scene. Prices are from the Sume pricing code (read 2026-10-10).

What is the endpoint chain?

Write the script as four short paragraphs, one per scene, so sentence segments line up with scenes.

  • POST /v1/tts-1.0/generate with text, segmentation: {mode: "sentence"} and timestamps: {words: true}. The result holds the audio URL, segment start times and total length.
  • POST /v1/images four times (google/nano-banana-2.1, 1K, 16:9), one per paragraph; run them together.
  • POST /v1/timeline-1.0/render with audio.url and four video[] slots whose start values are the first sentence start of each paragraph. Slot one must start at 0.
  • POST /v1/video-captions with words from the TTS result, so the captions follow the voice without a transcription pass.

What does it cost?

TTS bills by character at $0.0475 per 1,000, so a 900-character intro is under five cents. The rest is flat.

Newsletter video bill (read 2026-10-10)
StepBasisCost
Voice900 chars x $0.0475 per 1,000$0.0428
Four stills4 x $0.10 (1K)$0.40
Render1 started minute$0.10
Captions1 job, from TTS words$0.20
Total$0.7428

What are the limits?

Send only one of script_text, words, cues or segments to the caption job. Words from TTS are the right pick because the text and timing already match; speech-to-text would be an extra step and an extra chance for a wrong word. The Video captions page lists the fields.

Timeline accepts up to 200 slots and an audio spine of 1 to 1,800 seconds, so a long newsletter is possible but the video stops being a good fit around two to three minutes. Plan first: the free /v1/timeline-1.0/plan call returns the same validation as render without the charge, as shown in the plan-first training explainer.

The newsletter link or sign-up button cannot be clicked in a video. Put the link in the post, and use a short caption cue at the end of the spine with the reader's next step.

What do I do with a long issue?

Pick the one story with a number or a deadline in it and make that the video; leave the rest in the email. If you want a recurring series, a fixed four-scene structure keeps cost at the same $1.00 every week. A related pattern is the 60-second explainer chain, which swaps the stills for Wan clips and costs more.

Pre-publish checklist

Write the script so that each scene is two sentences or fewer. A scene with one long sentence holds one still for too long, and a scene with five sentences needs several stills.

Run the plan call first. It is free, checks every field the render would check, and tells you the output length before you pay for a started minute.

  • Read the four segment start times and confirm the first is 0.
  • Keep every URL on media.sume.com; generated stills and TTS audio already are.
  • Use a different Idempotency-Key per job so a retry never double-bills.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume