Beatoven maestro music and SFX API vs Sume Music Router
Beatoven's API makes music and sound effects from text. Sume's Music Router makes one track per call at a fixed $0.125 and has no SFX route. The differences.

Beatoven's API generates both music and sound effects from text prompts; Sume's Music Router generates music only, one audio file per call, at a fixed $0.125 per accepted generation. Pick Beatoven if you need effects from the same API, and Sume if music is one step in a video and voice pipeline you already run on Sume.
Beatoven facts are from its API page, read on 2026-10-02. That page shows no pricing, duration limits or rate limits, so none are compared here, and it carries a 2024 date, so check Beatoven's current docs.
What does Beatoven's API say it does?
The page says the API creates music and sound effects from text prompts, using an underlying model called maestro, trained on over 3 million ethically sourced music and sound effects tracks. It describes the output as custom, high-fidelity and cleared for commercial use, and quotes 10k+ tracks generated and over 100 developers using it.
Those are marketing figures from the vendor. They tell you the scope (music plus SFX, prompt controlled) but not the contract details you need for production, such as track length, formats or rate limits.
What does Sume's Music Router do?
POST /v1/music-router/generate takes a prompt of 1 to 5,000 characters and an optional image_url. Leaving out model or sending sume/music-auto lets Sume pick the engine (Lyria 3.5 today); the routable ids are sume/music-auto, lyria-3.5 and lyria-3-pro. GET /v1/music-router/models lists them.
Every router model charges the same fixed Music price per generation, $0.125, regardless of prompt length or image. The result is an audio artifact on media.sume.com, read from result.artifacts[] where type is audio. job.request.routed_model tells you which engine ran.
What are the real differences?
The table keeps to what each page states.
| Question | Beatoven (page) | Sume Music Router (docs) |
|---|---|---|
| Music from text | Yes | Yes |
| Sound effects from text | Yes | No route |
| Image input | Not stated | Optional image_url |
| Price | Not on the page | $0.125 per generation |
| Length control | Not stated | No duration field; ask in the prompt |
| Negative prompt | Not stated | Unsupported when non-empty; put exclusions in the prompt |
| Output | Not stated | Sume-hosted audio artifact, typically audio/mpeg |
What does a Sume call look like?
Exclusions go in the positive prompt, and length is asked for in words. Sending duration is rejected.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: beatoven-compare-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Calm corporate bed, 96 BPM, G major. Felt piano, soft pads, brushed kick. A 30-second track, no vocals, no spoken word."
}'What if I need sound effects too?
Sume has no SFX endpoint, so a pipeline that needs door slams or whooshes must source them elsewhere and join them with the music. Sume's timeline audio concat joins up to 20 Sume-hosted audio files gaplessly for $0.01 per job, which is how a bed and stingers become one file. You still have to host those files on Sume first by importing them.
If the one thing you want is one vendor for both music and effects, Beatoven's page describes exactly that. If you want music next to voice, avatar and render jobs on one wallet and one job API, use Sume.
How should I test the two before choosing?
Write one brief you would really use, for example a 30-second bed for a product clip, and run it through both. Judge four things: whether the track follows the tempo, instruments and mood you named, whether the length is close to what you asked for, how the vendor describes licensing for your use, and what the call cost. Beatoven's page does not list a price, so ask for one in writing before a comparison is meaningful.
On the Sume side, remember that the prompt is the only control. There is no seed, temperature, guidance or duration parameter, so the same prompt can return a different track, and a seven-part brief (emotion, genre, tempo as a number, key, two to four instruments, one named moment, era) gets closer than a mood word. Verify the generated audio instead of trusting the brief; the docs call these creative directions, not guaranteed output settings. Re-running costs another $0.125, so a handful of takes is cheap to budget.
Sources
Related posts
More in Comparisons
- Best TTS model right now: the leaderboard versus Sume's router
Eleven v4 leads the Artificial Analysis TTS board today. What that means if you generate speech through Sume, whose router serves Sonic models only.
- Cheapest 720p AI video per second: Veo, xAI, Sume
At 720p, Veo 3.1 lists $0.05 to $0.40 per second, xAI lists grok-imagine-video at $0.05, and Sume lists Grok Imagine Video 1.5 at $0.0125. Read 2026-10-01.
- Claude directory drops MCPB: remote vs local server, and Sume
Claude's directory no longer accepts MCPB desktop extensions. A remote HTTPS server needs no package; here is what that means for Sume.
- Cloudflare Stream clip API vs Sume video trim: start and end
Cloudflare Stream clips by posting a source UID with startTimeSeconds and endTimeSeconds. Sume video-trim takes start plus end or duration, returning a new MP4.
Written by Sume