Avatar video soundtrack: send prompt or audio_url, volume 0.05-0.4
Sume's Avatar Video package accepts a soundtrack with exactly one of prompt or audio_url and a volume from 0.05 to 0.4, default 0.15. What each costs and does.
The Avatar Video create request has an optional package.soundtrack object, and exactly one of prompt or audio_url is required inside it (not both, no mode field). prompt (1 to 5000 characters) generates a Music 1.0 bed and muxes it under the talking head; audio_url is a public HTTPS file you already have, mirrored into Sume media and muxed. volume is a linear level from 0.05 to 0.4, default 0.15. All of this is from the Avatar Video request schema in Sume's OpenAPI file; the music prompt follows the Music 1.0 docs.
What is the difference between prompt and audio_url?
The schema says the prompt path reserves a Music 1.0 estimate, while the audio_url path mirrors your file and muxes only, with no Music reservation and a mux cost of $0. Localhost and private URLs are rejected for audio_url, and a durable public URL is preferred. Create and preview check the URL shape; full accessibility is checked at generate-video, where a dead URL is a soft fail.
| Field | Rule | Notes |
|---|---|---|
prompt | 1 to 5000 characters | Music 1.0 generates the bed; exclusive with audio_url |
audio_url | Public HTTPS URL | Mirrored then muxed; no Music reservation |
volume | 0.05 to 0.4, default 0.15 | Linear bed level under speech |
Which should I use?
Use audio_url when you have licensed music already, or when you want the same bed across many videos. Use prompt when you want a new bed per video; write it like any music brief and end with "Instrumental, no vocals" so the bed does not compete with the speaker. Remember that Avatar 1.0 itself is English-only, so a soundtrack does not change the language limit.
What does a soundtrack block look like?
The block sits inside package in the Avatar Video create body; other required package fields are unchanged and are not shown.
{
"package": {
"soundtrack": {
"prompt": "Warm acoustic bed, 92 BPM, G major. Nylon guitar and soft shaker, a gentle lift at 0:15. Instrumental, no vocals.",
"volume": 0.12
}
}
}What if I want more control than a single volume?
Render the avatar video without a package soundtrack and mix in Timeline, where the bed has gain_db, fade_out_seconds and duck_db. The package field is the short path; Timeline is the control path.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar video preview: approve the first frame, then pick quality
Sume avatar video previews are tier-independent: approve stills, then render at standard, plus or max. What can change at generate-video, and what cannot.
- Avatar preview resource_status vs job_status: which field to poll
Avatar video previews return resource_status and job_status next to a legacy status field. Read resource_status for readiness and job_status for polling.
- Cancel an avatar video job: the 409 job_generation_already_started
You can cancel an avatar video job only before generation starts. After that Sume returns 409 job_generation_already_started and the job runs to completion.
- Create an AI avatar from profile traits: the props input
Avatar 1.0 can build a reusable avatar from structured traits, not a prompt or photo. The props input takes ethnicity, sex and age. When to use it.
Written by Sume