TTS word timestamps: timestamps.words and sentence segmentation

Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.

3 min readSume
All posts

To time captions to a generated voice, ask the TTS request for word timestamps: set timestamps: { words: true }. Add segmentation: { mode: "sentence" } to also split by sentence. Segmentation requires word timestamps; sending it without them is a validation error (segmentation_requires_word_timestamps).

Fields

From the Sume API request schema
FieldValuesDefault
timestamps.wordsbooleanoff
segmentation.modesentencenone
segmentation.boundary_lead_ms0 to 50070
segmentation.emit_audiobooleantrue
speedslow, normal, fastunset
language2 to 16 charactersunset
output_formatmp3, 44100 Hz, 128 kbps unless setmp3

A request

Send exactly one of transcript and transcript_source; both or neither is a validation error. The voice id must be in the library id shape.

import os, requests

r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
    headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
             "Idempotency-Key": "vo-timed-001"},
    json={
        "transcript": "Cut here. Then fade out slowly.",
        "voice": {"id": os.environ["VOICE_ID"]},
        "timestamps": {"words": True},
        "segmentation": {"mode": "sentence", "boundary_lead_ms": 70},
    })
print(r.status_code)
print(r.json())

Why use it

Without timings, captions for a synthetic voice would need a second pass: transcribe the audio you just made. That is another job, at one cent per audio minute, plus a chance the transcript differs from the script. Asking for timestamps at synthesis keeps text and time together.

Characters bill as usual at $0.0475 per 1,000: the 31-character line above reserves about $0.0015. For finished clips that already have speech, Video captions is the burn-in step at $0.20 a job.

Related posts

More in Developers

All Developers posts

Written by Sume