TTS word timestamps: timestamps.words and sentence segmentation
Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.

To time captions to a generated voice, ask the TTS request for word timestamps: set timestamps: { words: true }. Add segmentation: { mode: "sentence" } to also split by sentence. Segmentation requires word timestamps; sending it without them is a validation error (segmentation_requires_word_timestamps).
Fields
| Field | Values | Default |
|---|---|---|
timestamps.words | boolean | off |
segmentation.mode | sentence | none |
segmentation.boundary_lead_ms | 0 to 500 | 70 |
segmentation.emit_audio | boolean | true |
speed | slow, normal, fast | unset |
language | 2 to 16 characters | unset |
output_format | mp3, 44100 Hz, 128 kbps unless set | mp3 |
A request
Send exactly one of transcript and transcript_source; both or neither is a validation error. The voice id must be in the library id shape.
import os, requests
r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "vo-timed-001"},
json={
"transcript": "Cut here. Then fade out slowly.",
"voice": {"id": os.environ["VOICE_ID"]},
"timestamps": {"words": True},
"segmentation": {"mode": "sentence", "boundary_lead_ms": 70},
})
print(r.status_code)
print(r.json())Why use it
Without timings, captions for a synthetic voice would need a second pass: transcribe the audio you just made. That is another job, at one cent per audio minute, plus a chance the transcript differs from the script. Asking for timestamps at synthesis keeps text and time together.
Characters bill as usual at $0.0475 per 1,000: the 31-character line above reserves about $0.0015. For finished clips that already have speech, Video captions is the burn-in step at $0.20 a job.
Related posts
More in Developers
- Turn a roleplay debrief into an avatar feedback clip in Python
Take the written debrief from a roleplay or survey session and render it as a 16:9 Sume avatar clip with a retry-safe key, a 12 to 168 word check and polling.
- 12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4
A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.
- Vercel 800 s max duration: do you still need a Sume webhook?
Vercel Pro allows 800 s functions and a 30-minute beta. A Sume video job can still outlast one request, so use async or webhook mode and return in seconds.
- Verify x-sume-webhook-signature in Node: sume-v1 HMAC, raw body
A node:crypto verifier for Sume's sume-v1 signature that refuses an empty secret, checks the 5-minute window, and accepts either entry during a rotation.
Written by Sume