How do I add a listen-to-this-page audio version with TTS?

Turn each article into an audio file with one async TTS job per page: a 9,000-character article costs 43 cents on Sume. What it does not replace.

4 min readSume
All posts

To add a listen-to-this-page player, render the article text to audio once when the page is published, store the file next to the page, and point a plain audio element at it. On Sume that is one asynchronous TTS job per article: a 1,500-word page is about 9,000 characters and costs $0.43 at $0.0475 per 1,000 characters, billed once, not per listen. A hundred articles come to $43.00.

This is a convenience feature for readers who prefer to listen, and it helps people with low vision or reading difficulty. It does not replace semantic HTML, headings, alt text or a screen reader, which read the live page and which your readers already configure. Treat the audio file as an addition, never as the accessible version.

Why generate at publish time, not on click

Sume TTS is asynchronous. A request returns a job id immediately; the client polls the job or receives a signed webhook, and the finished audio is a hosted file. The optional sync mode only waits up to 30 seconds for the answer, and the service has no streaming output. A button that calls the API on click would make a reader wait for a job, so the better design is to run the job from your publish step and cache the file.

That also keeps the cost predictable. The job is billed from the character count of what you send, so you can compute the price of a page before you submit it, and cap a batch by multiplying.

Prepare the text, not the HTML

Send the article body as plain text. Strip navigation, cookie banners and code blocks, expand abbreviations your readers would hear badly, and spell out numbers the voice could misread. Spaces and punctuation count toward the billed length, so the text you send is the text you pay for.

Split long articles at section breaks. A request accepts up to 20,000 characters, and a job whose audio runs past 1,200 seconds fails without capturing a credit. At the 750 characters a minute implied by Cartesia's pricing page, that is about 15,000 characters per job, so a long read is two or three jobs joined with a Timeline audio concat at a flat $0.01 per join.

A publish-step script

The Node script below submits one job per article with an idempotency key built from the page slug, so a retried deploy never bills the same page twice. It prints the job id; poll /v1/jobs/{id}/status until terminal is true and read the audio artifact from /v1/jobs/{id}/result.

const body = {
  transcript: process.argv[2],
  voice: { id: process.env.VOICE_ID },
  language: "en",
  output_format: { container: "mp3", sample_rate: 44100, bit_rate: 128000 },
  mode: "async",
};

const res = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SUME_API_KEY}`,
    "Content-Type": "application/json",
    "Idempotency-Key": `listen-${process.env.PAGE_SLUG}`,
  },
  body: JSON.stringify(body),
});
const out = await res.json();
console.log((out.data ?? out).request_id);

Cost of the audio layer for a content site

The table prices three site sizes. The per-page figure assumes the 9,000-character article above; rate and rounding are as of the 2026-10-07 TTS 1.0 catalog entry. The vendor columns use each vendor's published per-million-character rate, applied to the same character count and without any rounding, to show the order of magnitude.

Audio for 9,000-character articles. Vendor rates read 2026-10-07; Sume rate from the TTS 1.0 catalog.
ServiceRate per 1M charactersPer page100 pages
Sume TTS 1.0$47.50$0.43$43.00
MAI-Voice-2.1$22$0.198$19.80
ElevenLabs Flash / Turbo$40$0.36$36.00
ElevenLabs v3$80$0.72$72.00

Choices that matter for readers

The 1-cent minimum per job also shapes how you batch. A 60-character page summary costs a cent on its own, so group short items such as summaries or captions into one request when the reader will hear them back to back, and keep one job per article for the long reads.

Finally, test the result on a real phone with the player you ship. Hosted files are plain mp3 by default, so the audio element needs no special library, and the file can sit behind your own CDN once you have copied it from the job result.

  • Pick mp3 at the default 44.1 kHz and 128 kbps for a page player; use wav only if you will edit or join the file.
  • Set language for every non-English article. A voice and language mismatch returns a 409 before any charge, which is cheaper than finding it in review.
  • Regenerate a page's audio only when its text changes, and key the job on the page slug plus a revision number.
  • Label the player so readers know the voice is synthetic.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume