AI song generator API: section markers in, lyrics out, $0.125
Sume Music Router writes songs from a 5,000-character prompt: put [0:00-0:30] section markers and a BPM in, read result.lyrics out. $0.125 each.

"AI song generator API" and "AI song maker API" are on Google's autocomplete today (read 2026-10-07). Sume's answer is the Music Router: POST /v1/music-router/generate with a prompt, and Lyria 3.5 as the default engine. The request has no seed, temperature, guidance or duration. Everything you control sits in the prompt, which is up to 5,000 characters.
A generation costs a fixed $0.125 (read 2026-10-07), whatever the prompt length.
Write a prompt that gives structure
The docs recommend a brief with seven axes: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and era or production. Add section markers when you want the song cut into parts that line up with a video, and say the length you want in words.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: song-001" \
-d '{"prompt": "Hushed neo-soul, 72 BPM, D minor. Rhodes through tape wow, brushed snare, muted trumpet. [0:00-0:20] Intro: Rhodes only. [0:20-0:50] Verse with the trumpet answering. [0:50-1:20] Chorus, full band. A 1-minute-20 track, 1998 production, dry and close."}'What comes back
Poll GET /v1/jobs/{id}/status until result_ready is true, then read GET /v1/jobs/{id}/result. The audio is in result.artifacts[] where type is audio (usually audio/mpeg on media.sume.com). When they are present, result.lyrics holds the model-reported lyrics or section map. The docs call those model-reported metadata, not an audio measurement, so listen to the track before you trust the map.
| Field | Rule |
|---|---|
prompt | 1 to 5,000 characters; put exclusions in the positive text |
negative_prompt | Non-empty value returns 400 negative_prompt_unsupported |
duration, duration_seconds | Rejected; write the length in the prompt |
image_url | Optional public HTTPS still for mood; same price |
| Instrumental | End the prompt with: Instrumental, no vocals. |
Get a sung track
The docs recommend ending an instrumental brief with "Instrumental, no vocals." For a song with a voice, leave that clause out and describe the vocal in the genre line. Add "no spoken word" only when the track sits under narration. If a prompt is rejected by policy, change the flagged content and keep the musical brief.
Sources
Related posts
More in Developers
- aiohttp web server: verify a Sume webhook with await request.read()
An aiohttp 3.14 web handler that verifies the Sume HMAC on await request.read(), then calls request.json(). Exits on an empty secret and handles webhook.test.
- Alt text for a 30-image gallery in one Sume Agent Completion
One Agent Completion call can take up to 30 images and return an alts array under an object schema. Python stdlib script with the cap, poll and limits.
- API key scopes for Sume: which key can call which endpoint family?
Sume API keys carry fixed scopes: formats:write, actions:read, agent_completions:write, account:read. Which scope each route needs, and why old keys get a 403.
- Arabic speech to text API: Sume STT with language_code ar
Transcribe Arabic audio with Sume STT: send language_code ar, check the reported language, and review the text. $0.01 per audio minute.
Written by Sume