Listen to This Article: Audio Player With Read-Along Highlighting
Build a read-along article player: ask Sume TTS for word timestamps and sentence segments, save the timings as JSON, and highlight text from currentTime.

A read-along player needs one thing the audio file alone does not have: when each sentence starts and ends. Sume's TTS returns that if you ask for it. Send timestamps.words: true and segmentation.mode: "sentence", store the timings next to the MP3, and move a highlight from the audio element's currentTime.
Facts below come from the Sume TTS request schema in the API reference and the repository, and from MDN's timeupdate page, read 2026-10-03.
What the job returns
Timings are in seconds from the start of the audio. Segments are gapless: each one ends exactly where the next begins, with a cut 70 ms after the last word of a sentence by default. With an MP3 output, which is the default at 44.1 kHz and 128 kbps, you get segment timings but no per-segment audio_url. For a player you only need the timings.
- Segmentation requires
timestamps.words: true; the API refuses it otherwise. - Audio that runs past 1,200 seconds fails with
tts_duration_exceededand no credit is captured, and one request holds at most 20,000 characters. Split a long article at paragraph breaks and add each earlier chunk'sduration_secondsto the next chunk's times. - The catalog price is $0.0475 per 1,000 characters, so a 9,000-character article costs about $0.43.
| Field | Shape | Use |
|---|---|---|
| audio_url | Mirrored audio file | The src of your audio element |
| duration_seconds | Number, present with word timestamps | Offset for the next chunk |
| words[] | word, start, end | Per-word highlight |
| segments[] | index, text, start, end, duration_seconds | Per-sentence highlight |
Generate the audio and the timing file
This submits the article, polls the job and writes article-audio.json. Set VOICE_AVATAR to an avatar handle whose voice is ready.
import json, os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
text = open("article.txt", encoding="utf-8").read()
body = {"transcript": text, "avatar_handle": os.environ["VOICE_AVATAR"],
"timestamps": {"words": True}, "segmentation": {"mode": "sentence"}}
r = requests.post(f"{API}/v1/tts-1.0/generate", json=body,
headers={**H, "Idempotency-Key": "article-audio-001"})
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H).json()["data"]
if s["terminal"]:
break
time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(job["result_url"], headers=H).json()["data"]["result"]
out = {"audio_url": res["audio_url"], "duration": res["duration_seconds"],
"sentences": [{"text": g["text"], "start": g["start"], "end": g["end"]}
for g in res["segments"]]}
json.dump(out, open("article-audio.json", "w"))Highlight from currentTime
Do not drive the highlight from timeupdate. MDN says the event fires between about 4 Hz and 66 Hz depending on system load, so a short sentence can come and go between two events. Read currentTime in requestAnimationFrame while the audio plays.
Render the read-along text from sentences[].text, not from your original HTML. The segments are cut where a token ends in ., ! or ?, so a name such as "Dr." can split a sentence, and indexes from your own splitter would then drift.
const a = document.querySelector("audio");
const box = document.querySelector("#readalong");
fetch("/article-audio.json").then((r) => r.json()).then(({ sentences }) => {
const els = sentences.map((s, i) => {
const p = Object.assign(document.createElement("span"), { textContent: s.text + " " });
p.onclick = () => { a.currentTime = s.start; a.play(); };
box.append(p);
return p;
});
let cur = -1, raf = 0;
const tick = () => {
const t = a.currentTime;
const i = sentences.findIndex((s) => t >= s.start && t < s.end);
if (i !== cur) { els[cur]?.classList.remove("on"); els[i]?.classList.add("on"); cur = i; }
raf = requestAnimationFrame(tick);
};
a.addEventListener("play", () => { cancelAnimationFrame(raf); tick(); });
a.addEventListener("pause", () => cancelAnimationFrame(raf));
});For finer highlighting, use words[] the same way. For cutting the voiceover into separate sentence files instead, see sentence clips with TTS segmentation.
Sources
Related posts
More in Developers
- LiteLLM /mcp-rest/tools/call: test Sume tools without an LLM
LiteLLM exposes REST routes to list and call MCP tools with no model in the loop. Use them to smoke-test Sume's read tools and gates before an agent sees them.
- Live AI avatar API: a Tavus conversation vs a Sume job
A live avatar API creates a room you join. Sume's Avatar API creates a job you poll. Field-by-field map of Tavus create conversation and Sume talking-video.
- LM Studio mcp.json: add Sume's hosted MCP server with an API key
LM Studio 0.3.17 and later accept remote MCP servers in mcp.json. The exact entry for https://mcp.sume.com/mcp with a Bearer header, plus the gates to know.
- Load-test your Sume webhook receiver with a signed burst
Fire 500 correctly signed job.completed requests at your own receiver before a bulk run does it for real. A runnable Python harness and the 10-second budget.
Written by Sume