Beatoven maestro music and SFX API vs Sume Music Router

Beatoven's API makes music and sound effects from text. Sume's Music Router makes one track per call at a fixed $0.125 and has no SFX route. The differences.

5 min readSume
All posts

Beatoven's API generates both music and sound effects from text prompts; Sume's Music Router generates music only, one audio file per call, at a fixed $0.125 per accepted generation. Pick Beatoven if you need effects from the same API, and Sume if music is one step in a video and voice pipeline you already run on Sume.

Beatoven facts are from its API page, read on 2026-10-02. That page shows no pricing, duration limits or rate limits, so none are compared here, and it carries a 2024 date, so check Beatoven's current docs.

What does Beatoven's API say it does?

The page says the API creates music and sound effects from text prompts, using an underlying model called maestro, trained on over 3 million ethically sourced music and sound effects tracks. It describes the output as custom, high-fidelity and cleared for commercial use, and quotes 10k+ tracks generated and over 100 developers using it.

Those are marketing figures from the vendor. They tell you the scope (music plus SFX, prompt controlled) but not the contract details you need for production, such as track length, formats or rate limits.

What does Sume's Music Router do?

POST /v1/music-router/generate takes a prompt of 1 to 5,000 characters and an optional image_url. Leaving out model or sending sume/music-auto lets Sume pick the engine (Lyria 3.5 today); the routable ids are sume/music-auto, lyria-3.5 and lyria-3-pro. GET /v1/music-router/models lists them.

Every router model charges the same fixed Music price per generation, $0.125, regardless of prompt length or image. The result is an audio artifact on media.sume.com, read from result.artifacts[] where type is audio. job.request.routed_model tells you which engine ran.

What are the real differences?

The table keeps to what each page states.

Beatoven API and Sume Music Router, read 2026-10-02
QuestionBeatoven (page)Sume Music Router (docs)
Music from textYesYes
Sound effects from textYesNo route
Image inputNot statedOptional image_url
PriceNot on the page$0.125 per generation
Length controlNot statedNo duration field; ask in the prompt
Negative promptNot statedUnsupported when non-empty; put exclusions in the prompt
OutputNot statedSume-hosted audio artifact, typically audio/mpeg

What does a Sume call look like?

Exclusions go in the positive prompt, and length is asked for in words. Sending duration is rejected.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: beatoven-compare-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Calm corporate bed, 96 BPM, G major. Felt piano, soft pads, brushed kick. A 30-second track, no vocals, no spoken word."
  }'

What if I need sound effects too?

Sume has no SFX endpoint, so a pipeline that needs door slams or whooshes must source them elsewhere and join them with the music. Sume's timeline audio concat joins up to 20 Sume-hosted audio files gaplessly for $0.01 per job, which is how a bed and stingers become one file. You still have to host those files on Sume first by importing them.

If the one thing you want is one vendor for both music and effects, Beatoven's page describes exactly that. If you want music next to voice, avatar and render jobs on one wallet and one job API, use Sume.

How should I test the two before choosing?

Write one brief you would really use, for example a 30-second bed for a product clip, and run it through both. Judge four things: whether the track follows the tempo, instruments and mood you named, whether the length is close to what you asked for, how the vendor describes licensing for your use, and what the call cost. Beatoven's page does not list a price, so ask for one in writing before a comparison is meaningful.

On the Sume side, remember that the prompt is the only control. There is no seed, temperature, guidance or duration parameter, so the same prompt can return a different track, and a seven-part brief (emotion, genre, tempo as a number, key, two to four instruments, one named moment, era) gets closer than a mood word. Verify the generated audio instead of trusting the brief; the docs call these creative directions, not guaranteed output settings. Re-running costs another $0.125, so a handful of takes is cheap to budget.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume