Text-to-speech API in Node: submit, poll and save the MP3

A Node 18+ fetch example for the Sume TTS Router: submit with an Idempotency-Key, poll status_url, read the audio artifact and save an MP3 to disk.

5 min readSume
All posts

To call a text-to-speech API from Node, POST to https://api.sume.com/v1/tts-router/generate with a model, a transcript and a voice selector, then poll the status_url from the response until terminal is true, then read the audio URL from result_url. Node 18 or newer has fetch built in, so the example below needs no packages. It writes intro.mp3 and prints the byte count.

The trend behind the question is Microsoft's 2026-10-01 launch of MAI-Voice-2.1 and its Flash tier. The October tracker lists a vendor claim of 150 ms end to end for Flash (read 2026-10-06). That is a real-time number. Sume TTS is a different shape: an asynchronous job that returns a finished file. Streaming TTS is a documented non-goal of the router, so this post is for narration, captions and ads, not for a live conversation.

The whole script

Set SUME_API_KEY and SUME_AVATAR_HANDLE first. The handle is any avatar whose voice status is ready; a voice.id UUID works in its place. The Idempotency-Key header makes a retried POST return the same job instead of a second paid one.

const headers = { Authorization: `Bearer ${process.env.SUME_API_KEY}`,
  "Content-Type": "application/json", "Idempotency-Key": "intro-line-v1" };
const sleep = (s) => new Promise((r) => setTimeout(r, s * 1000));
const get = async (url) => {
  const r = await fetch(url, { headers });
  if (!r.ok) throw new Error(`${url} ${r.status} ${await r.text()}`);
  return (await r.json()).data;
};
async function main() {
  const res = await fetch("https://api.sume.com/v1/tts-router/generate", {
    method: "POST", headers,
    body: JSON.stringify({ model: "sonic-3.6", language: "en",
      transcript: "Welcome to the demo. This line was generated by an API.",
      avatar_handle: process.env.SUME_AVATAR_HANDLE }),
  });
  if (!res.ok) throw new Error(`submit ${res.status} ${await res.text()}`);
  const job = (await res.json()).data;
  let state = job;
  while (!state.terminal) {
    await sleep(state.next_poll_after_seconds ?? 2);
    state = await get(job.status_url);
  }
  const art = (await get(job.result_url)).result.artifacts.find((a) => a.type === "audio");
  const bytes = Buffer.from(await (await fetch(art.url)).arrayBuffer());
  require("fs").writeFileSync("intro.mp3", bytes);
  console.log("saved intro.mp3", bytes.length, "bytes");
}
main().catch((e) => { console.error(e.message); process.exit(1); });

What the loop relies on

The loop reads next_poll_after_seconds from the status envelope instead of a fixed sleep, and stops on terminal, which is true for completed, failed and canceled jobs alike. A failed job has no result, so the result_url call returns 409 job_not_completed and the script throws with the body. That is the right place to read the error code. The shapes are in Jobs and results.

The default container is MP3 at 44100 Hz and 128 kbps. Add output_format: { container: "wav", sample_rate: 48000 } if the file goes into an editor. Add timestamps: { words: true } and the result carries words[] as well, which is the input for captions.

What it costs

Every catalog row on the router has the same list price: $0.0475 per 1,000 characters, the provider list rate with Sume's 1.25 margin applied, per the API pricing page (read 2026-10-06). Spaces and punctuation count. The line in the script above is 55 characters, so the raw price is about a quarter of a cent; the docs say Sume rounds up to the cent, so expect the job to bill $0.01. Microsoft's tracker price for MAI-Voice-2.1 is $22 per 1M characters and $15 for Flash, which works out to $0.022 and $0.015 per 1,000. Sume is higher per character. What you buy with the difference is an API that already carries a voice-language guard, word timings, sentence segments and signed webhooks for the same job.

Where each value comes from (read 2026-10-06)
StepCallRead this field
SubmitPOST /v1/tts-router/generatedata.status_url, data.result_url
WaitGET status_urldata.terminal, data.next_poll_after_seconds
FetchGET result_urldata.result.artifacts[] with type audio
DownloadGET artifact urlThe MP3 bytes

Failures to handle

Three failures are worth handling in code. An unknown model returns a 400 with a catalog_url; GET /v1/tts-router/models lists the ids that work. A balance that cannot cover the estimate returns 402 insufficient_credits before any job starts. A burst can return 429 rate_limited or queue_full; retry the same POST with the same Idempotency-Key.

If you would rather not poll, add a webhook_url and handle the terminal job.completed event instead. The Webhooks page has the signature scheme.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume