Word timestamps for an Eleven v4 voice: the router says 400, use STT
The Sume TTS Router rejects timestamps on eleven-* ids. Get word timings by running the finished narration through Sume STT for one cent a minute.

You cannot ask an eleven-* model on the Sume TTS Router for word timestamps: the router returns a 400 for timestamps, and for segmentation, on those rows. The workaround is to treat the finished narration as audio and send it to Sume STT, which always returns word timings and costs $0.01 per audio minute.
This matters for captions, karaoke-style highlights and cutting a voice-over to picture, because those all need to know when each word starts. The Sonic rows can produce timings in the same call, the Eleven rows cannot, so the pipeline differs by model.
What each path returns
The refusal list for Eleven rows comes from the router code: avatar_id, avatar_handle, transcript_source, generation_config, speed, timestamps, segmentation and pronunciation_dict_id. Everything else about timing has to be recovered after the fact.
| Path | Word timings | Extra cost |
|---|---|---|
| TTS Router, Sonic row, with timestamps requested | Returned with the audio | None beyond TTS |
| TTS Router, eleven-* row | Request is refused with a 400 | Not available |
| TTS Router eleven-* audio, then STT | Always returned in words[] | $0.01 per audio minute |
The two-step recipe
First generate the narration with your Eleven voice. Output is mp3 in the 44.1 kHz family, which is fine as an STT input. Then submit the resulting public HTTPS URL to POST /v1/stt-1.0/transcribe. The STT request takes audio_url, an optional language_code hint, an optional duration_seconds between 1 and 600, and an optional segmentation object with mode set to sentence if you also want sentence groups derived from the words.
Pass duration_seconds when you know it. If you omit it, Sume reserves one minute for usage, which is harmless for short lines but can understate the reservation for a long file. Audio over ten minutes needs to be split first; the timeline audio endpoint cuts a Sume-hosted file into up to 20 ranges for one cent per call.
import os, json, time, urllib.request
BASE = "https://api.sume.com"
HEAD = {
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
}
def call(method, path, body=None):
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(BASE + path, data=data, headers=HEAD, method=method)
with urllib.request.urlopen(req) as resp:
return json.load(resp)
body = {
"audio_url": "https://example.com/narration.mp3",
"duration_seconds": 90,
"segmentation": {"mode": "sentence"},
"mode": "async",
}
job = call("POST", "/v1/stt-1.0/transcribe", body)
print(json.dumps(job, indent=2)[:1500])Cost of getting timings this way
STT is $0.01 per audio minute in Sume's golden pricing fixtures. A 3-minute narration therefore adds $0.03, and a 10-minute narration, the longest a single request accepts, adds $0.10. The TTS charge is separate and depends on the model row; the per-1,000-character figures are in the router voice and price breakdown.
If you can choose the voice, the cheaper design is to use a Sonic row when you need timestamps in the same call, and an Eleven row when the voice itself is the reason for the choice.
Two cautions
STT recognises speech, so the returned words are what the recogniser heard, not your source script. Numbers, brand names and abbreviations can come back spelled differently, which matters if you align captions to the script. Align on timing, and keep your own text for display.
Sentence segmentation is derived from the word timings. If the provider returns no timed words, the request fails closed with a typed error instead of returning guessed segments, so handle that error rather than assuming words[] is never empty. The API reference describes the job states and polling, and the models overview lists the audio endpoints side by side.
Sources
Related posts
More in Developers
- xAI video API polling: pending, done, failed vs Sume job states
xAI's video API returns a request_id and three states. Here is how they map onto Sume's five /v1/videos statuses, with a runnable Python polling loop.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
Written by Sume