Luma Dream Machine add_audio endpoint vs Sume generate_audio
Luma lets you POST /generations/{id}/audio to add sound to a finished clip. Sume sets audio at submit with generate_audio; how to get the same result.

Luma's Dream Machine API can add sound to a clip you already made: POST /generations/{id}/audio with generation_type set to add_audio and an optional audio prompt. Sume has no call that attaches audio to a finished video job. You choose audio when you submit, with generate_audio, or you lay a soundtrack under finished clips on a timeline.
The Luma facts below come from the Add Audio reference page and the FAQ, both read on 2026-10-02. Luma's FAQ page says the current API docs are the Luma Agents API, so the endpoints on the older Dream Machine reference pages may not match what you call today. The Sume facts come from Video generation and Timeline 1.0.
What does the Luma add_audio endpoint take?
The endpoint is a second step on an existing generation. You pass the generation id in the path and a small JSON body. The only required body field is generation_type, fixed to add_audio. prompt and negative_prompt are optional text for the audio, and callback_url is an optional webhook.
It returns a generation object, so it is asynchronous like the video itself: state moves through queued, dreaming, completed or failed, and the assets object carries the video URL. The page does not state a price, a maximum audio length or a supported-model list, so none of those are claimed here.
| Field | Where | Required | Meaning |
|---|---|---|---|
| id | Path | Yes | Generation to add audio to |
| generation_type | Body | Yes | Always add_audio |
| prompt | Body | No | Text prompt for the audio |
| negative_prompt | Body | No | Negative prompt for the audio |
| callback_url | Body | No | URI called on state changes |
How does Sume handle audio on a video?
On Sume, audio is a property of the generation request. POST /v1/videos accepts generate_audio, a boolean that defaults to the model's own audio capability. Each model in GET /v1/videos/models reports a generate_audio flag, so you can see before you spend whether a model makes sound at all.
Some models cannot turn sound off. The Video Router page says Gemini Omni Flash 1.1 always has native synced audio and that generate_audio: false is rejected for it. Check the flag and the model page rather than assuming the field toggles everything.
curl "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY"
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Idempotency-Key: clip-audio-001" \
-H "Content-Type: application/json" \
-d '{"model":"seedance-2","prompt":"Rain on a tin roof, close shot","duration":5,"generate_audio":true}'What if the clip already exists and has no sound?
There are two honest routes. The first is to resubmit with a model that generates audio and generate_audio on; the finished clip is not modified, you get a new job and pay for it. The second is to keep the silent clip and add sound outside the generation.
For the second route, Timeline 1.0 assembles one audio spine plus ordered video slots into an MP4, and its optional soundtrack bed takes a URL, gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20. That is mixing, not generation: it does not write new sound from a text prompt for a given clip. To make the sound itself, create it with a music or speech job first and feed the hosted URL in.
How do the two designs compare?
Luma's design makes audio a follow-up you can decide after seeing the picture. Sume's design makes it a submit-time choice with a visible capability flag, and treats mixing as a separate step. Neither is wrong; the difference is where the decision sits and what a retry costs.
Budget for the Sume approach by pricing the whole job up front. Sume reserves the workspace balance on submit at provider list times 1.25, and usage.cost on the poll response is the billable amount, so a resubmit shows up as a second reservation.
| Question | Luma Dream Machine API | Sume |
|---|---|---|
| Add audio to a finished clip | POST /generations/{id}/audio | No call; resubmit or mix on a timeline |
| Audio prompt | prompt and negative_prompt | Part of the video prompt |
| Capability check | Not stated on the page | generate_audio per model in /v1/videos/models |
| Webhook | callback_url | callback_url on /v1/videos (HTTPS) |
Which should you pick?
If your workflow is picture first, sound later, and you already use Luma, the add_audio call is the shortest path there. If you want one job that returns a clip with sound and a price known at submit, set generate_audio and pick a model whose flag is true.
Whichever you use, poll the job rather than waiting on a single request, and keep the job id: see Jobs and results for the status vocabulary.
What should you test before relying on either route?
Run one short clip through the path you plan to use and listen to it. Audio quality, language and timing are properties of the model, and neither vendor page quoted here makes a promise about them, so a test on your own prompt is the only evidence that counts.
Then check three things in the response: that the job reached a terminal status, that the output URL plays with sound, and that the billed amount matches what you expected. On Sume that last number is usage.cost; on Luma it is the change in the credit balance. Record all three for the first run, and you will have a baseline for every later batch.
Sources
Related posts
More in Comparisons
- Luma Dream Machine upscale: 540p to 4K per generation vs Sume
Luma upscales an existing generation to 540p, 720p, 1080p or 4k with POST /generations/{id}/upscale. What Sume offers instead, and what its docs do not say.
- MAI-Voice-2 voice prompting from 5-60 s of audio vs Sume TTS
Microsoft MAI-Voice-2 prompts a voice from 5 to 60 seconds of reference audio. Sume TTS selects an existing voice and takes no reference clip.
- Advantage+ background generation for catalog ads: Meta's 2-3% claim
Meta says background generation for catalog ads lifted conversions 2-3%. You can turn it off; to control the scene yourself, edit the product shot on Sume.
- Meta Muse for small business: approval before spend vs Sume MCP
Meta's Muse for small business asks approval before publishing, sending or spending. Sume's MCP uses an idempotency key, which is not human approval.
Written by Sume