AI music adds vocals under narration: the Sume prompt fix
Sume Music has no negative prompt: a non-empty one returns 400. End the prompt with Instrumental, no vocals; add no spoken word. Each retry is $0.125.

If AI music keeps adding vocals under your narration, do not send a negative_prompt to Sume: a non-empty value returns HTTP 400 with public_reason=negative_prompt_unsupported. Put the exclusion in the positive prompt instead. End with "Instrumental, no vocals." and, only when the track sits under a voiceover, add "no spoken word." Each attempt is a $0.125 generation, so three tries cost $0.375.
What the docs say to write
The Music 1.0 docs list the request fields. prompt is 1 to 5,000 characters, and exclusions belong inside it. negative_prompt is accepted only as an omitted field or an empty string. The same docs advise ending a brief with one clause, "Instrumental, no vocals." and adding "no spoken word" only under narration, because a bed that contains speech competes with the voiceover.
| Wording | Where | Result |
|---|---|---|
| negative_prompt: "vocals" | Request field | 400 negative_prompt_unsupported |
| negative_prompt: "" | Request field | Accepted, ignored |
| Instrumental, no vocals. | End of the prompt | Recommended clause |
| ... no spoken word. | End of the prompt, under narration only | Recommended addition for beds |
A bed brief that keeps out of the voice
Write for the space the voice needs. Name two to four sparse instruments with texture, give a tempo as a number and a key and mode, and describe the arc: for a bed, that is usually flat. A bed with a clear melody in the 1 to 3 kHz range fights speech, so choose pads, low bass and a soft pulse. Length is in the prompt, such as "A 45-second track", because Music 1.0 does not take duration or duration_seconds.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: bed-under-vo-01" \
-d '{
"prompt": "Warm ambient bed, 70 BPM, A minor. Soft analog pad, low sine bass, a faint tape-hiss pulse, no melody. Flat arc, no build. A 45-second track. Instrumental, no vocals, no spoken word."
}'Why the vocal clause is not a guarantee
A prompt is guidance for a generative model, not a filter. Sume documents the exclusion as a recommended clause, not as a switch, so a track can still contain a vocal-like texture, a hummed pad or a breath. That is why the order matters: place the clause last, where it reads as the final instruction, and keep the rest of the brief free of words that suggest a singer.
The cost of testing is small and fixed. At $0.125 per generation, five variants of one brief are $0.625. Generate them together with different Idempotency-Key values, listen to all five, and keep the one that stays out of the voice's way. Reusing the same key would return the first job, not a new take.
Choosing the router or the direct route
The Music Router takes the same prompt rules and picks an engine for you; the direct Music 1.0 surface is the fixed one. Both bill $0.125 per generation. After a job finishes, check which engine actually ran before you compare takes, since a routed job can differ from the requested one. The track length still comes from the words in your prompt, so say "A 45-second track" every time and compare lengths in the result.
If vocals still appear
Four steps, in the order that costs least.
- Listen first. Provider
lyrics, when present, are model-reported metadata, not a measurement of what you hear. - Remove words that invite singing: "anthem", "choir", "soulful" and any lyric-like phrase in quotes.
- Change one thing per retry, and write down the job id, so you know which edit worked.
- Lower the bed under the voice anyway. In a Timeline render
soundtrack.duck_db(0 to 20) lowers it under speech, and a bit of leaked vocal is easier to hide at -12 dB.
Sources
Related posts
More in Media tools
- Audio detach refuses a 31-minute video: trim it first, $0.43 total
Audio detach rejects sources over 1,800 seconds. A 31-minute video needs trimming into pieces first: four trims, four detaches and 31 minutes of STT cost $0.43.
- Audio of a 15-second span: one detach range costs 1 cent, not 3
Need only the sound of seconds 40 to 55? Send one audio-detach job with a range for $0.01. Trimming first and then detaching costs $0.03.
- Background removal at $0.0225 per image, flat for any photo size
Sume's background removal costs $0.0225 per image and the catalog says the price does not vary by image size. 2,000 photos cost $45; 40,000 cost $900.
- Best of 8 music takes costs $1.00 on Sume, then a 1-cent cut
Music 1.0 and the Music Router charge a flat $0.125 per track and reject a duration field. Generate 8 takes for $1.00, then cut the best one with a $0.01 split.
Written by Sume