AI song generator API: section markers in, lyrics out, $0.125

Sume Music Router writes songs from a 5,000-character prompt: put [0:00-0:30] section markers and a BPM in, read result.lyrics out. $0.125 each.

5 min readSume
All posts

"AI song generator API" and "AI song maker API" are on Google's autocomplete today (read 2026-10-07). Sume's answer is the Music Router: POST /v1/music-router/generate with a prompt, and Lyria 3.5 as the default engine. The request has no seed, temperature, guidance or duration. Everything you control sits in the prompt, which is up to 5,000 characters.

A generation costs a fixed $0.125 (read 2026-10-07), whatever the prompt length.

Write a prompt that gives structure

The docs recommend a brief with seven axes: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and era or production. Add section markers when you want the song cut into parts that line up with a video, and say the length you want in words.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: song-001" \
  -d '{"prompt": "Hushed neo-soul, 72 BPM, D minor. Rhodes through tape wow, brushed snare, muted trumpet. [0:00-0:20] Intro: Rhodes only. [0:20-0:50] Verse with the trumpet answering. [0:50-1:20] Chorus, full band. A 1-minute-20 track, 1998 production, dry and close."}'

What comes back

Poll GET /v1/jobs/{id}/status until result_ready is true, then read GET /v1/jobs/{id}/result. The audio is in result.artifacts[] where type is audio (usually audio/mpeg on media.sume.com). When they are present, result.lyrics holds the model-reported lyrics or section map. The docs call those model-reported metadata, not an audio measurement, so listen to the track before you trust the map.

Music Router request rules, read 2026-10-07
FieldRule
prompt1 to 5,000 characters; put exclusions in the positive text
negative_promptNon-empty value returns 400 negative_prompt_unsupported
duration, duration_secondsRejected; write the length in the prompt
image_urlOptional public HTTPS still for mood; same price
InstrumentalEnd the prompt with: Instrumental, no vocals.

Get a sung track

The docs recommend ending an instrumental brief with "Instrumental, no vocals." For a song with a voice, leave that clause out and describe the vocal in the genre line. Add "no spoken word" only when the track sits under narration. If a prompt is rejected by policy, change the flagged content and keep the musical brief.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume