n8n text to speech: turn text into an audio file
Do text to speech in n8n with HTTP Request nodes: one submits the text, a Wait loop checks the job, and one downloads the audio as a binary file.

To do text to speech in n8n, send the text to a text-to-speech API from an HTTP Request node, wait until the job finishes, then download the audio with a second HTTP Request node set to return a file, which later nodes can use as binary data. With Sume, the first node POSTs the text and a voice to https://api.sume.com/v1/tts-1.0/generate, a Wait-and-check loop runs until the job ends, and the result's audio_url is the file to fetch.
Sume has no n8n node; this is a plain HTTPS call. Node settings follow n8n's docs for the HTTP Request node and the Wait node; Sume fields follow the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, and Jobs and results. All were read on 2026-09-28. The polling loop is the one built node by node in the speech-to-text version of this workflow.
How do I set up the text to speech request in n8n?
One HTTP Request node submits the job. Give it a generic Bearer auth credential that holds your Sume API key, and an Idempotency-Key header built from the item's own id, as in the speech-to-text workflow; what changes for speech is the body. Map the text from the incoming item with an expression, such as {{ $json.text }}.
- Using Fields Below builds the JSON for you. In a hand-written JSON body, a double quote or a line break in the text has to be escaped, or the body stops being valid JSON.
avatar_handleis the voice selector that fits a plain name/value pair. The other one,voice.id(a voice UUID or avoi_library id), sits inside avoiceobject, so send it from a Using JSON body.
| Setting | Value | Why |
|---|---|---|
| Method, URL | POST, https://api.sume.com/v1/tts-1.0/generate | Sume's TTS 1.0 route. |
| Send Body | JSON, Using Fields Below | Name/value pairs sent as a JSON body. |
transcript | {{ $json.text }} | 1–20,000 characters; spaces and punctuation count. |
avatar_handle | An avatar whose voice is ready | Selects the voice; the API publishes no voice list. |
language | ja, ko, or another code | Needed for non-English text; omitted, it defaults to English. |
How does the workflow wait for the audio?
The submit answers at once with data.status_url, before any audio exists. Poll it with the same loop as for speech to text: a short Wait node, an authenticated GET of the status URL, and an If node whose false output leads back to the Wait node until terminal is true, the way n8n builds a loop. A slow job is not a failed one, so don't resubmit it.
To skip polling, set a Wait node to On Webhook Call and send its $execution.resumeUrl as the job's webhook_url; n8n AI video workflow covers that pattern and how the per-execution URL meets Idempotency-Key. Sume calls back only on terminal events: completed, failed, or canceled.
How do I get the audio file into n8n?
Once terminal is true, check that {{ $json.data.sume_status }} is completed. A failed or canceled job has no result: its result_url answers 409 job_not_completed, and by default the HTTP Request node returns success only for a 2xx response. Then GET {{ $json.data.result_url }}, authenticated like the submit. The result sits under data.result, and in current code data.result.audio_url is the speech file, a public artifact on media.sume.com. A second HTTP Request node downloads it:
- Method
GET, URL{{ $json.data.result.audio_url }}. - Add Option → Response, then Response Format: File, with a field name such as
datain Put Output in Field. - Later nodes read the file from that field. Another HTTP Request node, for example, can send it on with the n8n Binary File body type, naming the field as its Input Data Field Name.
- The file is MP3 at 44,100 Hz and 128 kbps unless the request sets
output_format.
Is there a free text to speech API for n8n?
Sume's isn't free. Text to speech costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on each item's transcript characters, spaces and punctuation included, so a workflow that voices every new row bills per row. One request takes up to 20,000 characters, and speech longer than 1,200 seconds fails with tts_duration_exceeded without capturing credits. Rates are on API pricing.
Why was my n8n text to speech request refused?
Two refusals involve the voice, and both arrive at submit, before any job exists:
400withinvalid_voice_id:voice.idis neither a voice UUID nor avoi_library id. Copy it verbatim, or sendavatar_handleinstead.409 tts_voice_language_mismatch(current code): the library voice records a different language from the request's. No job or charge exists yet; to keep the voice, resend the same body andIdempotency-Keywithconfirm_language_mismatch: true.
Sources
- Sume API reference
- API reference
- Jobs and results
- API pricing
- n8n Docs: HTTP Request node (read 2026-09-28)
- n8n Docs: HTTP Request credentials (read 2026-09-28)
- n8n Docs: Wait node (read 2026-09-28)
- n8n Docs: If node (read 2026-09-28)
- n8n Docs: Looping (read 2026-09-28)
- n8n Docs: Expressions (read 2026-09-28)
Related posts
More in Integrations
- n8n upload to TikTok with the HTTP Request node
n8n has no TikTok node in its docs. Upload with HTTP Request nodes: query creator info, init a Direct Post with FILE_UPLOAD, then PUT the bytes.
- n8n: upload a file to Google Drive from a URL
Download the file with n8n's HTTP Request node set to return a File, then upload that binary with the Google Drive node's Upload operation into a folder.
- n8n YouTube Shorts automation: make a video, then upload it
Automate YouTube Shorts in n8n: start a vertical video with an HTTP Request node, download the finished MP4, and post it with the YouTube node's Upload.
- PowerShell Invoke-RestMethod POST JSON with a Bearer token
Convert a hashtable with ConvertTo-Json -Depth, then call Invoke-RestMethod -Method Post -ContentType 'application/json' with a Bearer header.
Written by Sume