Bedtime story audio: slow narrator and lullaby bed on Sume

Make a bedtime story track: slow TTS narration, an instrumental lullaby from Music, mixed with Timeline ducking, then an mp3. Steps and cost.

5 min readSume
All posts

Generate the story with Sume TTS at a slower speed, generate an instrumental lullaby with Music, mix them in a Timeline render with the lullaby as a ducked soundtrack over one still image, then detach the audio as an mp3. Four jobs, and the music, mix and detach steps cost about $0.74 for a six-minute story, before TTS.

Timeline outputs an MP4 and has no audio-only mode, which is why the last step extracts the audio.

How do I make the narration slow and calm?

TTS accepts generation_config.speed from 0.6 to 1.5 and a free-text emotion guide, plus volume from 0.5 to 2. Start with a speed a little under 1 and an emotion such as calm, then listen. Set language for any non-English story, since an omitted language defaults to English.

One TTS job synthesizes at most 1,200 seconds of audio, so a story over 20 minutes needs chapters, joined with timeline audio.

What music prompt works for a lullaby bed?

Write a brief with tempo, mode, two or three instruments and an arc, and close with "Instrumental, no vocals, no spoken word." Put exclusions in the positive prompt, because Music rejects a non-empty negative_prompt. A line like "slow lullaby, 60 BPM, music box and soft strings, constant and gentle" is a starting point. Steer length in words ("a 2-minute track"), since there is no duration field, and set loop true in Timeline if the story is longer than the track.

How do I mix and export it?

Timeline's audio spine is the narration; the soundtrack is the bed, with gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20. The render needs at least one video slot, so use one still image held for the whole story. The request below shows the shape; start the bed quiet (gain_db -20) and listen.

Then call audio detach on the finished MP4 with format mp3 (or wav) and a range. A whole track beyond 900 seconds needs a range.

import json

body = {
    "audio": {"url": "https://media.sume.com/artifacts/artf_demo/story.wav",
              "duration_seconds": 360},
    "soundtrack": {"url": "https://media.sume.com/artifacts/artf_demo/lullaby.mp3",
                   "gain_db": -20, "loop": True, "fade_out_seconds": 8, "duck_db": 10},
    "video": [{"source_url": "https://media.sume.com/artifacts/artf_demo/moon.png",
               "start": 0, "duration": 360}],
    "output": {"width": 1080, "height": 1080, "fps": 24},
}
print(json.dumps(body)[:80], "...")
minutes = -(-body["audio"]["duration_seconds"] // 60)
print("timeline", round(minutes * 0.10, 2), "music", 0.125, "detach", 0.01,
      "total", round(minutes * 0.10 + 0.125 + 0.01, 3))

What does it cost?

Rates read from the docs on 2026-10-02:

Costs for a six-minute story, excluding TTS (Sume docs, read 2026-10-02).
StepRateSix-minute story
Music generation$0.125 per generation$0.125
Timeline render$0.10 per output minute, rounded up$0.60
Audio detach$0.01 per job$0.01
Total$0.735

What should I check before sharing it?

Listen to the whole track once for a word swallowed by the bed, and lower gain_db or raise duck_db if so. If the story is for children on a platform, read that platform's own rules; they are not covered here.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume