AI meditation music generator: calm tracks, longer sessions
An AI meditation music generator makes calm instrumental tracks from a text brief. How to prompt one, join tracks into a session, and add a voice.

An AI meditation music generator makes calm instrumental tracks from a text brief: you describe a slow tempo, soft sustained instruments and no drums or vocals, and it returns a track. Single tracks run up to a few minutes, so for a 20- or 30-minute session you generate several in the same key and tempo and join them into one file.
On Sume the Music Router generates each track at $0.125 per audio generation, plus a 5.5% agent fee by default. Prompting comes from Music 1.0, joining from Timeline audio, and the voice step from the Sume API reference and Audio detach, read on 2026-09-29. Nothing here is a health or sleep claim.
How do I write a prompt for meditation music?
Use the seven-axis brief the Music docs suggest (the general template is in AI music prompt examples), with every axis pointed at stillness. The axes are creative directions, not guaranteed settings, so listen to each take:
- Emotion, precisely: "hushed, unhurried, warm" rather than "relaxing".
- Tempo as a number, kept slow, and the same key on every track you plan to join.
- Two to four instruments with texture: "soft analog pad", "low drone", "distant singing bowl".
- An arc with one gentle moment, e.g. "a second pad enters at 1:00".
- Exclusions in the prompt itself ("no percussion"), since a non-empty
negative_promptreturns 400, and the close "Instrumental, no vocals." - Length in the prompt ("a 3-minute track");
durationis rejected.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: calm-session-a-part-1" \
-d '{
"prompt": "Hushed, unhurried and warm ambient drone, 60 BPM, D major. Soft analog pad, low cello drone, distant singing bowl. A second pad enters at 1:00. A 3-minute track. No percussion. Instrumental, no vocals."
}'How do I make a 30-minute meditation track?
Join the tracks you keep with a Timeline audio concat. The join is sample-domain, with no silence at the seams, but it has no crossfade setting: each part starts where the last one ends. That is why matching key, tempo and instruments across the parts matters.
| Rule | What the docs say |
|---|---|
| Parts | 1 to 20, in order; each { url, source_in?, duration? } |
| Inputs | This workspace's media.sume.com audio, such as your generated tracks |
| Channel layout | All parts must share one, or audio_parts_channel_mismatch |
| Output length | Up to 1,800 s (30 minutes) |
| Output format | WAV (pcm_s16le) by default, or MP3 |
| Price | $0.01 per job |
How do I add a guided meditation voice over the music?
Voice the script with text to speech at a slow pace: generation_config.speed goes down to 0.6. Then mix it in a Timeline 1.0 render, with the voice as the audio spine and a music track as the soundtrack, and pull the mix back out as audio with audio detach; How to mix voice with background music walks through those steps. The render runs for its declared audio.duration_seconds, which that post sets to the voice's length, so the guided part lasts as long as the reading. For a long session:
- A Timeline render's audio runs 1 to 1,800 seconds.
- Audio detach outputs at most 900 seconds per job; a longer mix needs one
rangerequest per stretch. - For music-only minutes before or after the voice, join the guided mix and your music tracks with one concat. Parts with different channel layouts are refused (
audio_parts_channel_mismatch), so compare the detach result'schannelswith your tracks first.
How much does AI meditation music cost?
Each generation is a flat $0.125 per audio generation; each join or detach is one small job; the voice is billed per character and the render per started output minute. The sessions below are example assumptions.
| Example | Price |
|---|---|
| One calm track, 3 takes | $0.38 |
| Music-only session: 10 takes, 8 kept and joined | $1.26 |
| Guided track: 3 takes, a 3,000-character script, a 5-minute mix | $1.03 |
Sources
Related posts
More in Use cases
- AI meditation video generator: voice, calm loop, music bed
An AI guided meditation video is a slow narration over a calm visual loop and a soft music bed that dips under the voice. How to build one, and costs.
- AI menu generator: make the artwork, type the prices
An AI menu generator works best for the artwork: backgrounds, borders, and dish photos. Type dish names and prices yourself. Sizes, prompts, and cost.
- AI movie trailer generator: shots, narration, music, titles
An AI movie trailer is short generated shots, a narrator, music that builds and burned-in title cards, cut together. How to make one, with costs.
- AI product photo retouching: fix defects, keep the product
AI product photo retouching: send the real photo to an image-edit model, list the defects to fix and what must not change, then check every file.
Written by Sume