Keep a text-to-speech pipeline portable: job ids, webhooks, files
How to build a TTS and music pipeline so a model or vendor change touches one function: idempotency keys, signed webhooks, polling, stored files, metadata.

A portable audio pipeline keeps four things outside the vendor: your scripts, your settings, your generated files and a record of which model made each file. Everything else is a function you can swap. Vendor moves make this concrete: TechCrunch reported on 2026-09-30 a $22 billion valuation for ElevenLabs, and Suno announced with v6 that older models will be retired, so a provider's model list can change at short notice.
Sume's job API has the pieces for this habit built in; the sections below show which ones to use.
Which fields should you store per job?
Store a row per generated file. The first three columns come from your own code; the last two come from the job.
| Field | Source | Why you keep it |
|---|---|---|
| Script text and revision | Your system | Regenerate on any engine |
| Request body | Your system | Reproduce the call |
| Idempotency-Key | Your system | Safe retries without double billing |
| job id and model | Sume job | Trace a file to an engine |
| Audio URL, downloaded copy | Sume result | The file survives a vendor change |
How do idempotency keys and webhooks help?
Send an Idempotency-Key on every submit, and reuse the same key only for the same payload, so a network retry never bills twice. Ask for mode: webhook with a public HTTPS webhook_url, and verify the signed terminal event. Sume sends terminal events only, so there is no progress feed to build on; keep polling status_url as a backup for missed deliveries, as the docs advise.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scene-12-v3" \
-d '{"transcript":"Scene twelve narration, final text.",
"avatar_handle":"@speaker",
"metadata":{"script_revision":"r7","scene":12}}'What does the metadata field do?
metadata is caller data stored with the Sume job request and not sent to the provider. Put your script revision and scene number there so a job can be matched to your records from either side. Do not put secrets or personal data in it. Keep the keys short and stable, such as script_revision and scene, so that exports from different weeks line up and a script can be re-rendered on another engine without hunting through old notes.
What should I do?
Wrap generation in one function that takes text and settings and returns a file and a record. Download every result when its job completes. Then a vendor change is a change to that function and a re-test, not a rebuild. See Jobs and results, Webhooks and the API reference.
Sources
Related posts
More in Developers
- Kimi Code CLI YOLO and AFK modes with paid Sume tools over MCP
Kimi auto-approves MCP tool calls in YOLO and AFK modes. What that means for a Sume server with write access, and the limits to set first.
- Kimi mcp test fails on Sume: URL, auth and scope checklist
If kimi mcp test fails or shows no Sume tools, check the endpoint URL, run kimi mcp auth, then separate a read-only grant from a broken connection.
- Kling 4.0 launch-day checklist for an API integration
Kling 4.0 is due in October. A short checklist and a Python script that flags the day a new Kling id appears in the Sume video catalog.
- Kling motion control with avatar_id instead of image_url
Sume's Kling 3.0 Motion Control takes image_url or avatar_id/avatar_handle, never both. A ready avatar resolves server-side to its identity still.
Written by Sume