Azure word boundaries in ms vs Sume TTS timestamps in seconds
Azure writes word timings as AudioOffset and Duration in milliseconds, in a separate file. Sume returns words[] with start and end seconds on the job result.

With Azure batch synthesis you turn on wordBoundaryEnabled or sentenceBoundaryEnabled and get a [nnnn].word.json or .sentence.json file in the ZIP, where each word has Text, AudioOffset and Duration in milliseconds. On Sume TTS 1.0 you send timestamps.words: true and the completed job result includes words[] with start and end in seconds. Convert units before you share code.
Azure's format is from its batch synthesis page, read 2026-10-01; Sume's from the API reference.
How do the two timing shapes compare?
Different file layout, different unit, same purpose.
| Item | Azure batch synthesis | Sume TTS 1.0 |
|---|---|---|
| Switch | wordBoundaryEnabled: true | timestamps.words: true |
| Where | Separate .word.json in the results ZIP | words[] on the completed job result |
| Fields | Text, AudioOffset, Duration | start, end per word |
| Unit | Milliseconds | Seconds |
| Sentences | sentenceBoundaryEnabled file | segmentation.mode: sentence |
What does Sume do for sentences?
Add segmentation: { "mode": "sentence" } with timestamps.words: true. The result has gapless segments[], where each segment ends exactly where the next starts, using a 70 ms post-word boundary by default that you can set from 0 to 500. With a wav or raw container each segment also gets its own audio URL; with mp3 you get timings only.
How do I convert an Azure word file?
Divide by 1,000. An Azure word with AudioOffset 200 and Duration 350 starts at 0.2 seconds and, if you want an end, ends at 0.55. Azure gives a duration, Sume gives an end, so compute end = (offset + duration) / 1000.
What does the request look like?
Ask for wav so the sentence slices come back as audio.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "The rainbow has seven colors.",
"avatar_handle": "@narrator",
"output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence" }
}'Sources
Related posts
More in Developers
- Azure Long Audio API retires April 2027: long audio on Sume
Azure says the Long Audio API retires April 1, 2027 in favor of batch synthesis. Sume TTS 1.0 handles long text as async jobs capped at 1,200 seconds of audio.
- BFL 24 concurrent requests, 6 for flux-kontext-max, vs Sume
BFL caps concurrent requests at 24 (6 for flux-kontext-max). Sume treats concurrency as a dispatch limit: extra jobs queue until you hit 429 queue_full.
- BFL 402 and 429 retry rules, and the same split on Sume
BFL raises on 402 and backs off on 429. Sume splits the same way: 402 insufficient_credits is a stop, 429 queue_full or rate_limited means wait and retry.
- BFL polling_url on api.bfl.ai vs Sume status_url
BFL says to always poll the polling_url it returns. Sume's job envelope carries status_url and result_url for the same reason: follow them, do not build URLs.
Written by Sume