Word timestamps for an Eleven v4 voice: the router says 400, use STT

The Sume TTS Router rejects timestamps on eleven-* ids. Get word timings by running the finished narration through Sume STT for one cent a minute.

4 min readSume
All posts

You cannot ask an eleven-* model on the Sume TTS Router for word timestamps: the router returns a 400 for timestamps, and for segmentation, on those rows. The workaround is to treat the finished narration as audio and send it to Sume STT, which always returns word timings and costs $0.01 per audio minute.

This matters for captions, karaoke-style highlights and cutting a voice-over to picture, because those all need to know when each word starts. The Sonic rows can produce timings in the same call, the Eleven rows cannot, so the pipeline differs by model.

What each path returns

The refusal list for Eleven rows comes from the router code: avatar_id, avatar_handle, transcript_source, generation_config, speed, timestamps, segmentation and pronunciation_dict_id. Everything else about timing has to be recovered after the fact.

Where word timings come from, per path, from the Sume router code and the STT schema, read 2026-10-11
PathWord timingsExtra cost
TTS Router, Sonic row, with timestamps requestedReturned with the audioNone beyond TTS
TTS Router, eleven-* rowRequest is refused with a 400Not available
TTS Router eleven-* audio, then STTAlways returned in words[]$0.01 per audio minute

The two-step recipe

First generate the narration with your Eleven voice. Output is mp3 in the 44.1 kHz family, which is fine as an STT input. Then submit the resulting public HTTPS URL to POST /v1/stt-1.0/transcribe. The STT request takes audio_url, an optional language_code hint, an optional duration_seconds between 1 and 600, and an optional segmentation object with mode set to sentence if you also want sentence groups derived from the words.

Pass duration_seconds when you know it. If you omit it, Sume reserves one minute for usage, which is harmless for short lines but can understate the reservation for a long file. Audio over ten minutes needs to be split first; the timeline audio endpoint cuts a Sume-hosted file into up to 20 ranges for one cent per call.

import os, json, time, urllib.request

BASE = "https://api.sume.com"
HEAD = {
    "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
    "Content-Type": "application/json",
}

def call(method, path, body=None):
    data = json.dumps(body).encode() if body is not None else None
    req = urllib.request.Request(BASE + path, data=data, headers=HEAD, method=method)
    with urllib.request.urlopen(req) as resp:
        return json.load(resp)

body = {
    "audio_url": "https://example.com/narration.mp3",
    "duration_seconds": 90,
    "segmentation": {"mode": "sentence"},
    "mode": "async",
}
job = call("POST", "/v1/stt-1.0/transcribe", body)
print(json.dumps(job, indent=2)[:1500])

Cost of getting timings this way

STT is $0.01 per audio minute in Sume's golden pricing fixtures. A 3-minute narration therefore adds $0.03, and a 10-minute narration, the longest a single request accepts, adds $0.10. The TTS charge is separate and depends on the model row; the per-1,000-character figures are in the router voice and price breakdown.

If you can choose the voice, the cheaper design is to use a Sonic row when you need timestamps in the same call, and an Eleven row when the voice itself is the reason for the choice.

Two cautions

STT recognises speech, so the returned words are what the recogniser heard, not your source script. Numbers, brand names and abbreviations can come back spelled differently, which matters if you align captions to the script. Align on timing, and keep your own text for display.

Sentence segmentation is derived from the word timings. If the provider returns no timed words, the request fails closed with a typed error instead of returning guessed segments, so handle that error rather than assuming words[] is never empty. The API reference describes the job states and polling, and the models overview lists the audio endpoints side by side.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume