ACE-Step 1.5 open-source music vs Lyria on Sume's Music Router
ACE-Step 1.5 is MIT-licensed, makes up to 10-minute songs and runs in under 4 GB of VRAM. How it compares with Lyria 3.5 via Sume's Music Router.

ACE-Step 1.5 is an MIT-licensed music model you run yourself, with songs from 10 seconds to 10 minutes, LoRA training and selective regeneration, while Sume offers music as a hosted job on Lyria 3.5 with one fixed price per generation. Run ACE-Step if you want to fine-tune and iterate for free on your own GPU; use Sume's Music Router if you want a finished mp3 per request and no infrastructure.
ACE-Step facts come from its repository and Lyria facts from Google's music generation docs, both read on 2026-10-02. Sume facts come from the Music Router docs.
What can ACE-Step 1.5 do?
The README lists generation from 10 seconds to 10 minutes (600 seconds), multi-language lyrics in 50 or more languages, a vocal-to-BGM conversion feature, cover generation, stem separation and repaint, which regenerates a selected part. It describes one-click LoRA training in Gradio, with 8 songs taking about an hour on an RTX 3090.
Hardware tiers run from under 4 GB of VRAM for basic operation to 24 GB or more for the largest model, with CUDA, ROCm, Intel XPU, Apple Silicon (MLX) and CPU paths. Speed is stated as under 2 seconds per full song on an A100 and under 10 seconds on an RTX 3090. The README positions quality between Suno v4.5 and Suno v5. That is the project's own claim, so listen before you rely on it.
What does Sume's Music Router give you?
POST /v1/music-router/generate takes a prompt of up to 5,000 characters and an optional image_url. Omit model or send sume/music-auto and Sume picks the engine, which is Lyria 3.5 today. Explicit ids from the catalog are lyria-3.5 and lyria-3-pro. Every model charges the same fixed Music price per audio generation, $0.125 at the catalog price, regardless of prompt length.
There is no duration field, no negative prompt, no stems and no repaint. You steer length and structure inside the prompt, for example with section markers such as [0:00-0:30] Intro:. Google's docs say Lyria 3.5 is single-turn: you cannot edit a result with a follow-up prompt, and identical prompts give different results.
| Feature | ACE-Step 1.5 | Sume Music Router |
|---|---|---|
| Length | 10 seconds to 10 minutes, set by you | Steered in the prompt; no duration field |
| Editing | Repaint, covers, stems | New take only; no iterative editing |
| Customization | LoRA training on your own songs | Prompt only; no fine-tuning |
| Cost | Your hardware | $0.125 per generation, whichever engine |
| Output | Files on your machine | Sume-hosted audio artifact (typically audio/mpeg) |
What does a Sume request look like?
Put the structure in the prompt, close with the exclusion, and send an idempotency key so a retry never double-charges.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: score-001" \
-d '{
"model": "sume/music-auto",
"prompt": "[0:00-0:15] Intro: warm lo-fi hip hop, 84 BPM, C minor, dusty Rhodes. [0:15-0:45] Groove: brushed boom-bap drums, muted trumpet answer. A 45-second track. Instrumental, no vocals."
}'
Read the audio artifact from the result:
curl https://api.sume.com/v1/jobs/$JOB_ID/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/$JOB_ID/result \
-H "Authorization: Bearer $SUME_API_KEY"
Which should you use?
Use ACE-Step for experiments that need control: repainting a bridge, building a LoRA on your own catalog, or generating hundreds of drafts without per-track billing. You carry the setup, the GPU time and the rights review of whatever you train on.
Use Sume when the music is one deliverable among others, such as a bed under a video. You can drop the audio into a Timeline 1.0 soundtrack with gain_db, loop, fade_out_seconds and duck_db instead of editing it by hand. Either way, check the rights terms for commercial use and the labeling rules for AI-made music before publishing.
A practical way to decide is to run one prompt on both. Write a 45-second brief with tempo, key and an arc, generate three takes locally and three through the router, and compare them on the same speakers. Count the minutes you spent on setup as well as the minutes of audio you got. If the local takes are better and you will make many of them, the setup pays back. If you need one clean take per video, the hosted price is hard to beat.
Sources
Related posts
More in Comparisons
- Is Firefly output 'commercially safe'? Firefly vs partner models
Adobe says outputs from Firefly AI models are safe for commercial use and trained on licensed and public domain content. Partner models are listed separately.
- AI disclosure API fields: TikTok, Instagram, YouTube and Snap
The four field names that mark a video as AI-made: TikTok is_aigc, Instagram is_ai_generated, YouTube containsSyntheticMedia, Snap ai_content_source.
- AI image text: what each vendor claims and how to test it
OpenAI admits text placement can fail, Google promises legible stylized text. Four vendor pages, what they actually claim, and a test harness to run on Sume.
- AI video cost per minute: Veo 3.1 vs the Sume catalog
One minute of generated footage costs $4.80 to $24 on Veo 3.1 at list prices and $0.75 to $15 on the Sume video catalog. Per-second rates, dated 2026-10-01.
Written by Sume