Keep a text-to-speech pipeline portable: job ids, webhooks, files

How to build a TTS and music pipeline so a model or vendor change touches one function: idempotency keys, signed webhooks, polling, stored files, metadata.

5 min readSume
All posts

A portable audio pipeline keeps four things outside the vendor: your scripts, your settings, your generated files and a record of which model made each file. Everything else is a function you can swap. Vendor moves make this concrete: TechCrunch reported on 2026-09-30 a $22 billion valuation for ElevenLabs, and Suno announced with v6 that older models will be retired, so a provider's model list can change at short notice.

Sume's job API has the pieces for this habit built in; the sections below show which ones to use.

Which fields should you store per job?

Store a row per generated file. The first three columns come from your own code; the last two come from the job.

Per-job record for a portable pipeline; Sume fields from the Jobs and results docs, read 2026-10-02.
FieldSourceWhy you keep it
Script text and revisionYour systemRegenerate on any engine
Request bodyYour systemReproduce the call
Idempotency-KeyYour systemSafe retries without double billing
job id and modelSume jobTrace a file to an engine
Audio URL, downloaded copySume resultThe file survives a vendor change

How do idempotency keys and webhooks help?

Send an Idempotency-Key on every submit, and reuse the same key only for the same payload, so a network retry never bills twice. Ask for mode: webhook with a public HTTPS webhook_url, and verify the signed terminal event. Sume sends terminal events only, so there is no progress feed to build on; keep polling status_url as a backup for missed deliveries, as the docs advise.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: scene-12-v3" \
  -d '{"transcript":"Scene twelve narration, final text.",
       "avatar_handle":"@speaker",
       "metadata":{"script_revision":"r7","scene":12}}'

What does the metadata field do?

metadata is caller data stored with the Sume job request and not sent to the provider. Put your script revision and scene number there so a job can be matched to your records from either side. Do not put secrets or personal data in it. Keep the keys short and stable, such as script_revision and scene, so that exports from different weeks line up and a script can be re-rendered on another engine without hunting through old notes.

What should I do?

Wrap generation in one function that takes text and settings and returns a file and a record. Download every result when its job completes. Then a vendor change is a change to that function and a re-test, not a rebuild. See Jobs and results, Webhooks and the API reference.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume