How do I add a listen-to-this-page audio version with TTS?
Turn each article into an audio file with one async TTS job per page: a 9,000-character article costs 43 cents on Sume. What it does not replace.

To add a listen-to-this-page player, render the article text to audio once when the page is published, store the file next to the page, and point a plain audio element at it. On Sume that is one asynchronous TTS job per article: a 1,500-word page is about 9,000 characters and costs $0.43 at $0.0475 per 1,000 characters, billed once, not per listen. A hundred articles come to $43.00.
This is a convenience feature for readers who prefer to listen, and it helps people with low vision or reading difficulty. It does not replace semantic HTML, headings, alt text or a screen reader, which read the live page and which your readers already configure. Treat the audio file as an addition, never as the accessible version.
Why generate at publish time, not on click
Sume TTS is asynchronous. A request returns a job id immediately; the client polls the job or receives a signed webhook, and the finished audio is a hosted file. The optional sync mode only waits up to 30 seconds for the answer, and the service has no streaming output. A button that calls the API on click would make a reader wait for a job, so the better design is to run the job from your publish step and cache the file.
That also keeps the cost predictable. The job is billed from the character count of what you send, so you can compute the price of a page before you submit it, and cap a batch by multiplying.
Prepare the text, not the HTML
Send the article body as plain text. Strip navigation, cookie banners and code blocks, expand abbreviations your readers would hear badly, and spell out numbers the voice could misread. Spaces and punctuation count toward the billed length, so the text you send is the text you pay for.
Split long articles at section breaks. A request accepts up to 20,000 characters, and a job whose audio runs past 1,200 seconds fails without capturing a credit. At the 750 characters a minute implied by Cartesia's pricing page, that is about 15,000 characters per job, so a long read is two or three jobs joined with a Timeline audio concat at a flat $0.01 per join.
A publish-step script
The Node script below submits one job per article with an idempotency key built from the page slug, so a retried deploy never bills the same page twice. It prints the job id; poll /v1/jobs/{id}/status until terminal is true and read the audio artifact from /v1/jobs/{id}/result.
const body = {
transcript: process.argv[2],
voice: { id: process.env.VOICE_ID },
language: "en",
output_format: { container: "mp3", sample_rate: 44100, bit_rate: 128000 },
mode: "async",
};
const res = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": `listen-${process.env.PAGE_SLUG}`,
},
body: JSON.stringify(body),
});
const out = await res.json();
console.log((out.data ?? out).request_id);Cost of the audio layer for a content site
The table prices three site sizes. The per-page figure assumes the 9,000-character article above; rate and rounding are as of the 2026-10-07 TTS 1.0 catalog entry. The vendor columns use each vendor's published per-million-character rate, applied to the same character count and without any rounding, to show the order of magnitude.
| Service | Rate per 1M characters | Per page | 100 pages |
|---|---|---|---|
| Sume TTS 1.0 | $47.50 | $0.43 | $43.00 |
| MAI-Voice-2.1 | $22 | $0.198 | $19.80 |
| ElevenLabs Flash / Turbo | $40 | $0.36 | $36.00 |
| ElevenLabs v3 | $80 | $0.72 | $72.00 |
Choices that matter for readers
The 1-cent minimum per job also shapes how you batch. A 60-character page summary costs a cent on its own, so group short items such as summaries or captions into one request when the reader will hear them back to back, and keep one job per article for the long reads.
Finally, test the result on a real phone with the player you ship. Hosted files are plain mp3 by default, so the audio element needs no special library, and the file can sit behind your own CDN once you have copied it from the job result.
- Pick mp3 at the default 44.1 kHz and 128 kbps for a page player; use wav only if you will edit or join the file.
- Set
languagefor every non-English article. A voice and language mismatch returns a 409 before any charge, which is cheaper than finding it in review. - Regenerate a page's audio only when its text changes, and key the job on the page slug plus a revision number.
- Label the player so readers know the voice is synthetic.
Sources
Related posts
More in Developers
- Mandarin Chinese speech to text API: Sume STT language_code zh
Transcribe Mandarin audio with Sume STT using language_code zh, then check the result and timings. $0.01 per audio minute and no accuracy claim without a test.
- Migrate a real-time avatar prototype to Sume async jobs: what changes
Moving from a live avatar session to Sume means replacing a stream with submit, poll and fetch. The code changes, the UX changes, and a Node example that runs.
- Model an AI generation job as a state machine in your database
A schema and update rule for tracking Sume jobs: five statuses, sticky terminal states, a separate webhook delivery column, and the idempotency key on the row.
- Pin the model id in an ad test: sume/auto follows the catalog
sume/auto is a pure function of the request plus the catalog version, so two ad arms made weeks apart can land on different models. Pin an explicit id in tests.
Written by Sume