Documentary underscore with AI music: a 3-minute brief and cost
How to brief an instrumental underscore for a 3-minute documentary with narration: one Lyria request, a ducked bed in Timeline, and the full audio cost.

For a documentary with narration, ask for one instrumental underscore of about three minutes, keep its arc gentle so it never competes with the voice, and mix it under the spoken track with a duck. On Sume that is one request to POST /v1/music-router/generate at $0.125, then a Timeline render at $0.10 per started output minute, so a 180-second film's audio and assembly come to $0.425.
The part that takes thought is the brief. Narration carries the information, so the music has to hold a mood without a melody that fights the sentences. The prompt below asks for sparse, slow material that grows only slightly, and it ends with the instrumental clause the Music docs recommend.
The brief to send
Put the arc in the text. The Music 1.0 page says the prompt controls song length and structure, and that a request such as a 2-minute track or section markers like [0:00-0:30] Intro works. Here the markers carve the three minutes into a patient opening, a middle that thickens a little, and a long release.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: documentary-narration-underscore-ai-musi-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Reflective documentary underscore, 66 BPM, D dorian. Felt piano, bowed double bass, soft low strings, a distant clarinet answering at 1:10. [0:00-0:40] Sparse piano, room tone. [0:40-2:20] Strings enter, slow swell. [2:20-3:00] Strands fall away to one piano line. Dry and intimate, 2010s indie-film production. No spoken word. Instrumental, no vocals. A 3-minute track."
}'Why each axis is set that way
| Axis | Choice |
|---|---|
| Emotion | Reflective and unforced, nothing triumphant |
| Genre | Indie-film underscore, piano and strings |
| Tempo | 66 BPM, slow enough to sit under speech |
| Key and mode | D dorian, minor colour without heaviness |
| Instruments | Felt piano, bowed bass, one clarinet answer |
| Arc | One named moment: strings enter at 0:40, release at 2:20 |
| Era and production | Dry, close, modern indie-film |
Getting the length right
The router has no duration field. Sume rejects duration and duration_seconds, so the three minutes live in the prompt, as above. Treat the result as approximate and check the real length of the returned MP3 before you build on it.
In the Timeline render, set audio.duration_seconds to your narration length (180 for this film) and pass the MP3 as soundtrack.url. If the track is a little short, soundtrack.loop repeats it; if it is a little long, the spine decides the output length. Add fade_out_seconds (up to 10) for a clean ending, and duck_db between 0 and 20 to lower the bed whenever the narration speaks. Ducking needs a real audio spine, so a silent spine will be refused with duck_requires_audio_spine.
What it costs
Narration, not the bed, sets the length of the film, so the render is billed on the three-minute spine. The plan endpoint, POST /v1/timeline-1.0/plan, is unbilled and returns billable_minutes and an estimated cost before you commit.
| Line | Rate | Amount |
|---|---|---|
| Music generation x 1 | $0.125 each | $0.125 |
| Timeline render, 180 s = 3 started minutes | $0.10 per started minute | $0.30 |
| Total for the audio and assembly | $0.425 |
Check before you publish
A documentary lives or dies on the voice, so the test is simple: play the first minute with the bed ducked and ask whether any sentence needed a second listen. If one did, regenerate with a sparser arc or raise duck_db.
Keep the job id, prompt and routed engine (job.request.routed_model) with the cut so the track can be reproduced or replaced later.
- Read the returned audio before approval. Provider lyrics are model-reported metadata, not a measurement of the track.
- Keep exclusions in the positive prompt. A non-empty
negative_promptis rejected with HTTP 400. - Use the same Idempotency-Key on a retry so a timeout does not make a second paid generation.
- Confirm usage terms for generated music with Sume before a commercial release; the docs read for this post state a price, not a license.
Sources
Related posts
More in Media tools
- Does Kling motion control make a photo talk? Where sound comes from
Kling 3.0 motion control copies movement from a reference video onto a still and takes no audio input. For speech from audio, use Fabric or H3 Max on Sume.
- End a vertical Short on a fade: Timeline fade_out_seconds 0.5
Timeline output.fade_in_seconds and fade_out_seconds run 0 to 5 s, and their sum cannot exceed the output length. Request, limits, and error code.
- Esports highlight reel music: a 45-second hype bed with a drop
Brief a 45-second esports highlight with a drop at 0:20: one Lyria generation at $0.125, a Timeline render and a duck under game commentary.
- Fade 0.25 s between two vertical clips: Timeline transition limits
Timeline transitions run up to 1 second and half of the shorter neighbor, never on the first slot, and no more than 8 chained. Request and error codes.
Written by Sume