n8n text to speech: turn text into an audio file

Do text to speech in n8n with HTTP Request nodes: one submits the text, a Wait loop checks the job, and one downloads the audio as a binary file.

6 min readSume
All posts

To do text to speech in n8n, send the text to a text-to-speech API from an HTTP Request node, wait until the job finishes, then download the audio with a second HTTP Request node set to return a file, which later nodes can use as binary data. With Sume, the first node POSTs the text and a voice to https://api.sume.com/v1/tts-1.0/generate, a Wait-and-check loop runs until the job ends, and the result's audio_url is the file to fetch.

Sume has no n8n node; this is a plain HTTPS call. Node settings follow n8n's docs for the HTTP Request node and the Wait node; Sume fields follow the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, and Jobs and results. All were read on 2026-09-28. The polling loop is the one built node by node in the speech-to-text version of this workflow.

How do I set up the text to speech request in n8n?

One HTTP Request node submits the job. Give it a generic Bearer auth credential that holds your Sume API key, and an Idempotency-Key header built from the item's own id, as in the speech-to-text workflow; what changes for speech is the body. Map the text from the incoming item with an expression, such as {{ $json.text }}.

  • Using Fields Below builds the JSON for you. In a hand-written JSON body, a double quote or a line break in the text has to be escaped, or the body stops being valid JSON.
  • avatar_handle is the voice selector that fits a plain name/value pair. The other one, voice.id (a voice UUID or a voi_ library id), sits inside a voice object, so send it from a Using JSON body.
From n8n's HTTP Request node docs and the TTS schema in the Sume API reference, read 2026-09-28.
SettingValueWhy
Method, URLPOST, https://api.sume.com/v1/tts-1.0/generateSume's TTS 1.0 route.
Send BodyJSON, Using Fields BelowName/value pairs sent as a JSON body.
transcript{{ $json.text }}1–20,000 characters; spaces and punctuation count.
avatar_handleAn avatar whose voice is readySelects the voice; the API publishes no voice list.
languageja, ko, or another codeNeeded for non-English text; omitted, it defaults to English.

How does the workflow wait for the audio?

The submit answers at once with data.status_url, before any audio exists. Poll it with the same loop as for speech to text: a short Wait node, an authenticated GET of the status URL, and an If node whose false output leads back to the Wait node until terminal is true, the way n8n builds a loop. A slow job is not a failed one, so don't resubmit it.

To skip polling, set a Wait node to On Webhook Call and send its $execution.resumeUrl as the job's webhook_url; n8n AI video workflow covers that pattern and how the per-execution URL meets Idempotency-Key. Sume calls back only on terminal events: completed, failed, or canceled.

How do I get the audio file into n8n?

Once terminal is true, check that {{ $json.data.sume_status }} is completed. A failed or canceled job has no result: its result_url answers 409 job_not_completed, and by default the HTTP Request node returns success only for a 2xx response. Then GET {{ $json.data.result_url }}, authenticated like the submit. The result sits under data.result, and in current code data.result.audio_url is the speech file, a public artifact on media.sume.com. A second HTTP Request node downloads it:

  • Method GET, URL {{ $json.data.result.audio_url }}.
  • Add Option → Response, then Response Format: File, with a field name such as data in Put Output in Field.
  • Later nodes read the file from that field. Another HTTP Request node, for example, can send it on with the n8n Binary File body type, naming the field as its Input Data Field Name.
  • The file is MP3 at 44,100 Hz and 128 kbps unless the request sets output_format.

Is there a free text to speech API for n8n?

Sume's isn't free. Text to speech costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on each item's transcript characters, spaces and punctuation included, so a workflow that voices every new row bills per row. One request takes up to 20,000 characters, and speech longer than 1,200 seconds fails with tts_duration_exceeded without capturing credits. Rates are on API pricing.

Why was my n8n text to speech request refused?

Two refusals involve the voice, and both arrive at submit, before any job exists:

  • 400 with invalid_voice_id: voice.id is neither a voice UUID nor a voi_ library id. Copy it verbatim, or send avatar_handle instead.
  • 409 tts_voice_language_mismatch (current code): the library voice records a different language from the request's. No job or charge exists yet; to keep the voice, resend the same body and Idempotency-Key with confirm_language_mismatch: true.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume