Text-to-speech API in Node: submit, poll and save the MP3
A Node 18+ fetch example for the Sume TTS Router: submit with an Idempotency-Key, poll status_url, read the audio artifact and save an MP3 to disk.

To call a text-to-speech API from Node, POST to https://api.sume.com/v1/tts-router/generate with a model, a transcript and a voice selector, then poll the status_url from the response until terminal is true, then read the audio URL from result_url. Node 18 or newer has fetch built in, so the example below needs no packages. It writes intro.mp3 and prints the byte count.
The trend behind the question is Microsoft's 2026-10-01 launch of MAI-Voice-2.1 and its Flash tier. The October tracker lists a vendor claim of 150 ms end to end for Flash (read 2026-10-06). That is a real-time number. Sume TTS is a different shape: an asynchronous job that returns a finished file. Streaming TTS is a documented non-goal of the router, so this post is for narration, captions and ads, not for a live conversation.
The whole script
Set SUME_API_KEY and SUME_AVATAR_HANDLE first. The handle is any avatar whose voice status is ready; a voice.id UUID works in its place. The Idempotency-Key header makes a retried POST return the same job instead of a second paid one.
const headers = { Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json", "Idempotency-Key": "intro-line-v1" };
const sleep = (s) => new Promise((r) => setTimeout(r, s * 1000));
const get = async (url) => {
const r = await fetch(url, { headers });
if (!r.ok) throw new Error(`${url} ${r.status} ${await r.text()}`);
return (await r.json()).data;
};
async function main() {
const res = await fetch("https://api.sume.com/v1/tts-router/generate", {
method: "POST", headers,
body: JSON.stringify({ model: "sonic-3.6", language: "en",
transcript: "Welcome to the demo. This line was generated by an API.",
avatar_handle: process.env.SUME_AVATAR_HANDLE }),
});
if (!res.ok) throw new Error(`submit ${res.status} ${await res.text()}`);
const job = (await res.json()).data;
let state = job;
while (!state.terminal) {
await sleep(state.next_poll_after_seconds ?? 2);
state = await get(job.status_url);
}
const art = (await get(job.result_url)).result.artifacts.find((a) => a.type === "audio");
const bytes = Buffer.from(await (await fetch(art.url)).arrayBuffer());
require("fs").writeFileSync("intro.mp3", bytes);
console.log("saved intro.mp3", bytes.length, "bytes");
}
main().catch((e) => { console.error(e.message); process.exit(1); });What the loop relies on
The loop reads next_poll_after_seconds from the status envelope instead of a fixed sleep, and stops on terminal, which is true for completed, failed and canceled jobs alike. A failed job has no result, so the result_url call returns 409 job_not_completed and the script throws with the body. That is the right place to read the error code. The shapes are in Jobs and results.
The default container is MP3 at 44100 Hz and 128 kbps. Add output_format: { container: "wav", sample_rate: 48000 } if the file goes into an editor. Add timestamps: { words: true } and the result carries words[] as well, which is the input for captions.
What it costs
Every catalog row on the router has the same list price: $0.0475 per 1,000 characters, the provider list rate with Sume's 1.25 margin applied, per the API pricing page (read 2026-10-06). Spaces and punctuation count. The line in the script above is 55 characters, so the raw price is about a quarter of a cent; the docs say Sume rounds up to the cent, so expect the job to bill $0.01. Microsoft's tracker price for MAI-Voice-2.1 is $22 per 1M characters and $15 for Flash, which works out to $0.022 and $0.015 per 1,000. Sume is higher per character. What you buy with the difference is an API that already carries a voice-language guard, word timings, sentence segments and signed webhooks for the same job.
| Step | Call | Read this field |
|---|---|---|
| Submit | POST /v1/tts-router/generate | data.status_url, data.result_url |
| Wait | GET status_url | data.terminal, data.next_poll_after_seconds |
| Fetch | GET result_url | data.result.artifacts[] with type audio |
| Download | GET artifact url | The MP3 bytes |
Failures to handle
Three failures are worth handling in code. An unknown model returns a 400 with a catalog_url; GET /v1/tts-router/models lists the ids that work. A balance that cannot cover the estimate returns 402 insufficient_credits before any job starts. A burst can return 429 rate_limited or queue_full; retry the same POST with the same Idempotency-Key.
If you would rather not poll, add a webhook_url and handle the terminal job.completed event instead. The Webhooks page has the signature scheme.
Sources
Related posts
More in Developers
- tts_sentence_selection_invalid 422 on Sume TTS: what triggers it
Sume TTS returns 422 tts_sentence_selection_invalid for gaps, repeated jobs, unfinished jobs and partial coverage. Each cause and its fix.
- tts_source_integrity_mismatch 422: job differs from accepted script
verify-spine returns 422 tts_source_integrity_mismatch when a finished TTS job's text no longer matches the accepted script. What it checks and how to recover.
- tts_source_not_found 404 on Sume TTS: revision, sentence or job
A 404 tts_source_not_found from the Sume script-source API means the revision, a sentence id or a selected job is not visible to this key or thread.
- tts_source_revision_mismatch 409: stale expected_script_revision_id
Sume returns 409 tts_source_revision_mismatch when the accepted script changed under you or a run is frozen. How to re-read the revision and retry.
Written by Sume