How do I make a spoken wake-up alarm with an AI voice and music?

A wake-up alarm is a short voice greeting joined to a music file: 2 cents of TTS and a 1-cent concat per day, with one $0.125 music track reused all week.

4 min readSume
All posts

Build a spoken wake-up alarm by generating a short voice greeting with TTS, generating one gentle music track with Music Router, and joining the two into one audio file with a Timeline audio concat. A 220-character greeting is $0.02 on Sume and the join is $0.01, so five different mornings cost $0.15 for the voice and the joins, plus $0.125 once for the track that all five reuse. Alarm apps play files, so the finished product is a plain wav or mp3 the user picks as a sound.

The voice plays first and the music follows it. Sume's documented audio tools join files end to end; the only documented way to lay music under a voice is a Timeline render, which outputs an MP4. For an alarm that has to be an audio file, sequence the pieces instead of layering them.

Write the greeting for the ear

Keep the greeting under 250 characters, which is about 20 seconds at the 750 to 800 characters a minute that Cartesia's pricing page gives for Sonic speech (one credit per character, the same metering Sume TTS 1.0 uses). Use short sentences, name the day, and end on a cue that the music is about to start: 'Good morning. It is Tuesday. Your first meeting is at ten. Here we go.'

At $0.0475 per 1,000 characters, 220 characters is $0.01045, which the catalog rounds up to $0.02. The 1-cent minimum per job means very short greetings still cost a cent or two, so batching five mornings into one request would be cheaper by the character but would give you one long file to split. Five small jobs are simpler.

Make the track once

Music Router rejects a duration field, so the length goes in the prompt. Ask for 'a 30-second gentle wake-up piece, 76 BPM, G major, soft piano and light strings, building slowly, instrumental, no vocals'. Every Music Router model charges the same $0.125 per generation, and you can pin an engine from the router catalog or leave the default sume/music-auto. Generate two or three candidates and keep the one that does not make you flinch at 6 a.m.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wake-up-bed-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "A 30-second gentle wake-up piece, 76 BPM, G major. Soft piano and light strings, building slowly. Instrumental, no vocals."
  }'

Join voice and track

Timeline audio concat takes up to 20 ordered parts, each a media.sume.com audio URL in your workspace, and returns one gapless file with no re-synthesis. Both parts must share a channel layout; if the job fails with audio_parts_channel_mismatch, check the layout of each file before you retry. Choose wav output while you are iterating; mp3 is smaller, but it adds encoder padding at each edge, which matters little for an alarm.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wake-up-tuesday-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/tuesday-greeting.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/wake-up-bed.mp3" }
    ],
    "output": { "format": "mp3" }
  }'

A week of mornings, priced

Small jobs show the rounding more than large ones. A 220-character greeting is 1.045 cents of characters but bills as 2 cents, nearly double. On MAI-Voice-2.1 at $22 per million characters the same greeting would be about 0.48 cents before any rounding on their side, and on ElevenLabs Flash at $40 per million about 0.88 cents, using the rates on each vendor's own page read 2026-10-07. For an app that sends thousands of tiny messages a day, that per-job rounding is the line to model; for five mornings a week it is pennies.

If you ship this as a feature, generate the week's greetings in one batch on Sunday night, store the files, and let the alarm app pick the file for the day. That keeps the user's wake-up time independent of any job's queue time, which is the safer design for anything that has to be ready at a fixed hour.

  • Reuse the same track and change only the greeting, so the sound is recognizable.
  • Test the file at the loudness of a phone speaker before you ship it; an alarm that is too quiet is a failure.
  • Keep greetings free of personal data you would not want read aloud on a lock screen.
Five spoken wake-up files on Sume, rates as of 2026-10-07 (TTS 1.0, Music Router and Timeline audio catalog entries); vendor figures read 2026-10-07.
ItemUnitPriceCountTotal
Greeting, 220 charactersper job (rounded up)$0.025$0.10
Music track, 30 secondsper generation$0.1251$0.125
Concat of greeting plus trackflat per job$0.015$0.05
Week total$0.275

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume