Halloween narrator voiceover with spooky music for a video
Make a Halloween story video audio track: a slow narrator from Sume TTS, an instrumental horror bed from the music router, mixed with ducking in Timeline 1.0.

To score a Halloween story video, generate the narration with Sume text-to-speech, generate an instrumental bed with the music router, and mix them in one Timeline 1.0 render with the bed ducked under the voice. Three calls, with two prices you can add up: $0.0475 per 1,000 characters of narration and $0.125 per accepted music generation.
Step 1: the narration
Write 400 to 800 characters per scene so a take stays short. Set language explicitly, since it defaults to English if omitted. Slow the voice down with generation_config.speed (range 0.6 to 1.5) and add a mood with generation_config.emotion, a free text field up to 64 characters, for example "hushed, ominous". Whether a given voice follows that hint closely is something to listen for, not assume.
Step 2: the music bed
The music router prompt takes 1 to 5,000 characters. There is no duration field, so say the length in the prompt. A prompt built from the seven axes in the music docs works well: emotion, genre, BPM, key, instruments with texture, an arc, and era. Close with "Instrumental, no vocals."
Foreboding, slow-burn cinematic horror. 62 BPM, D minor.
Bowed cello drone, detuned music box, distant low piano,
sparse reverse cymbal swells. A 2-minute track.
Stays quiet, then one sudden string stab at 1:30.
Analog 1970s horror score production. Instrumental, no vocals.Step 3: mix it
Pass the narration as the audio spine and the music as soundtrack in Timeline 1.0. Set soundtrack.gain_db, duck_db (0 to 20) to lower the bed while the voice speaks, fade_out_seconds (up to 10) so the music does not stop dead, and loop if the bed is shorter than the story. Ducking needs a real voice spine; it is rejected with a silent track. Render is $0.10 per output minute, and POST /v1/timeline-1.0/plan is unbilled, so check the plan first. Details: Timeline 1.0.
| Step | Price | Example |
|---|---|---|
| Narration | $0.0475 per 1,000 characters | 3,000 characters = $0.1425 |
| Music bed | $0.125 per accepted generation | one bed = $0.125 |
| Render | $0.10 per output minute | 1 minute = $0.10 |
Limits
The music router rejects a non-empty negative_prompt, so describe what you want rather than what to avoid. Generated music is not guaranteed to hit your timestamp, so listen to the stab before you publish. See looping a bed under a long video.
Related posts
More in Use cases
- Higgsfield's Seedance outage on Sept 30: slow vs failed jobs on Sume
Higgsfield said Seedance 2.0 failed more and 2.5 ran slow on Sept 30, now fixed. On Sume, a slow job is queued or processing; rerun only after failed.
- Check 200 holiday clips before upload with a free probe-only inspect
A video-inspect call with frames false is a probe with no stills and no charge. Use it as a pre-upload audio gate, then pull stills only for clips you doubt.
- 300 holiday UGC clips on one plan: pace around queue_full 429s
A Pro workspace holds 4 processing and 20 queued jobs. Submit 300 avatar clips, treat queued as normal, and retry queue_full with the same idempotency key.
- Hour-long podcast video to clips: the 1800-second source cap
Sume's trim, detach and inspect tools read sources up to 1800 seconds, so a 60-minute episode must be split before import. Limits and a safe split plan.
Written by Sume