Listen to This Article: Audio Player With Read-Along Highlighting

Build a read-along article player: ask Sume TTS for word timestamps and sentence segments, save the timings as JSON, and highlight text from currentTime.

5 min readSume
All posts

A read-along player needs one thing the audio file alone does not have: when each sentence starts and ends. Sume's TTS returns that if you ask for it. Send timestamps.words: true and segmentation.mode: "sentence", store the timings next to the MP3, and move a highlight from the audio element's currentTime.

Facts below come from the Sume TTS request schema in the API reference and the repository, and from MDN's timeupdate page, read 2026-10-03.

What the job returns

Timings are in seconds from the start of the audio. Segments are gapless: each one ends exactly where the next begins, with a cut 70 ms after the last word of a sentence by default. With an MP3 output, which is the default at 44.1 kHz and 128 kbps, you get segment timings but no per-segment audio_url. For a player you only need the timings.

  • Segmentation requires timestamps.words: true; the API refuses it otherwise.
  • Audio that runs past 1,200 seconds fails with tts_duration_exceeded and no credit is captured, and one request holds at most 20,000 characters. Split a long article at paragraph breaks and add each earlier chunk's duration_seconds to the next chunk's times.
  • The catalog price is $0.0475 per 1,000 characters, so a 9,000-character article costs about $0.43.
TTS result fields for a read-along, read 2026-10-03
FieldShapeUse
audio_urlMirrored audio fileThe src of your audio element
duration_secondsNumber, present with word timestampsOffset for the next chunk
words[]word, start, endPer-word highlight
segments[]index, text, start, end, duration_secondsPer-sentence highlight

Generate the audio and the timing file

This submits the article, polls the job and writes article-audio.json. Set VOICE_AVATAR to an avatar handle whose voice is ready.

import json, os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
text = open("article.txt", encoding="utf-8").read()
body = {"transcript": text, "avatar_handle": os.environ["VOICE_AVATAR"],
        "timestamps": {"words": True}, "segmentation": {"mode": "sentence"}}
r = requests.post(f"{API}/v1/tts-1.0/generate", json=body,
                  headers={**H, "Idempotency-Key": "article-audio-001"})
r.raise_for_status()
job = r.json()["data"]
while True:
    s = requests.get(job["status_url"], headers=H).json()["data"]
    if s["terminal"]:
        break
    time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(job["result_url"], headers=H).json()["data"]["result"]
out = {"audio_url": res["audio_url"], "duration": res["duration_seconds"],
       "sentences": [{"text": g["text"], "start": g["start"], "end": g["end"]}
                     for g in res["segments"]]}
json.dump(out, open("article-audio.json", "w"))

Highlight from currentTime

Do not drive the highlight from timeupdate. MDN says the event fires between about 4 Hz and 66 Hz depending on system load, so a short sentence can come and go between two events. Read currentTime in requestAnimationFrame while the audio plays.

Render the read-along text from sentences[].text, not from your original HTML. The segments are cut where a token ends in ., ! or ?, so a name such as "Dr." can split a sentence, and indexes from your own splitter would then drift.

const a = document.querySelector("audio");
const box = document.querySelector("#readalong");
fetch("/article-audio.json").then((r) => r.json()).then(({ sentences }) => {
  const els = sentences.map((s, i) => {
    const p = Object.assign(document.createElement("span"), { textContent: s.text + " " });
    p.onclick = () => { a.currentTime = s.start; a.play(); };
    box.append(p);
    return p;
  });
  let cur = -1, raf = 0;
  const tick = () => {
    const t = a.currentTime;
    const i = sentences.findIndex((s) => t >= s.start && t < s.end);
    if (i !== cur) { els[cur]?.classList.remove("on"); els[i]?.classList.add("on"); cur = i; }
    raf = requestAnimationFrame(tick);
  };
  a.addEventListener("play", () => { cancelAnimationFrame(raf); tick(); });
  a.addEventListener("pause", () => cancelAnimationFrame(raf));
});

For finer highlighting, use words[] the same way. For cutting the voiceover into separate sentence files instead, see sentence clips with TTS segmentation.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume