Music under narration: add "no spoken word" to the Lyria prompt

Instrumental, no vocals ends a Sume music brief. When the track sits under a voiceover, add no spoken word too, so no voice competes with yours.

5 min readSume
All posts

Sume's music brief guide ends every instrumental prompt with "Instrumental, no vocals". When the track will sit under a voiceover, add "no spoken word" as well. Lyria does not accept a negative prompt through Sume (a non-empty negative_prompt returns 400 negative_prompt_unsupported), so the exclusions must be written in the positive prompt.

Why the extra clause

"No vocals" tells the model not to sing. A track can still carry chopped vocal samples, breathy phrases or a spoken intro, particularly in genres where those are part of the sound. Under a narrator, any voice in the bed competes with the one you paid for. "No spoken word" closes that gap.

Sume's operations guide on music brief variance says to add it only when narration is present, which keeps briefs without a voiceover shorter.

What the full brief looks like

The brief has seven axes: emotion, genre, BPM, key and mode, two to four textured instruments, an arc with a named moment, and era or production. The last sentence is the exclusion.

Each generation costs a fixed $0.125, so a missed clause costs a re-run.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: bed-narration-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Reflective documentary underscore, 72 BPM, D minor. Felt piano, bowed double bass, a distant cello swell at 0:25. A 1-minute track. Instrumental, no vocals, no spoken word."
  }'

Why you cannot use a negative prompt

Some music tools expose a negative field. Sume's router rejects a non-empty negative_prompt with negative_prompt_unsupported, because the Lyria path takes none. Put "no spoken word" in the prompt itself. See the negative prompt post for the details.

Lyria also does not have a duration field on Sume. Say the length in words, such as "a 1-minute track".

Check the result

Listen to the first 30 seconds and the end, where intros and outros are likely to carry vocal textures. result.lyrics on the job is model-reported metadata, and the vendor notes that results are non-deterministic and vary between runs. If a voice slips through, re-run with the clause moved to the front of the prompt, or pick another take.

Under the narration, duck the bed with duck_db so the words stay in front. Ducking needs the voice file as the Timeline spine. The duck walkthrough has the render.

Keep the arc simple under a voice

A voiceover needs room, so ask for a bed with a gentle arc and one named moment rather than a busy build. Two to four textured instruments is the guidance, and sparse ones leave space for the words. If the voice still fights the track, raise duck_db in Timeline before you regenerate.

Regenerating costs $0.125 a time; a ducking change is free to try on the render.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume