ACE-Step 1.5 open-source music vs Lyria on Sume's Music Router

ACE-Step 1.5 is MIT-licensed, makes up to 10-minute songs and runs in under 4 GB of VRAM. How it compares with Lyria 3.5 via Sume's Music Router.

5 min readSume
All posts

ACE-Step 1.5 is an MIT-licensed music model you run yourself, with songs from 10 seconds to 10 minutes, LoRA training and selective regeneration, while Sume offers music as a hosted job on Lyria 3.5 with one fixed price per generation. Run ACE-Step if you want to fine-tune and iterate for free on your own GPU; use Sume's Music Router if you want a finished mp3 per request and no infrastructure.

ACE-Step facts come from its repository and Lyria facts from Google's music generation docs, both read on 2026-10-02. Sume facts come from the Music Router docs.

What can ACE-Step 1.5 do?

The README lists generation from 10 seconds to 10 minutes (600 seconds), multi-language lyrics in 50 or more languages, a vocal-to-BGM conversion feature, cover generation, stem separation and repaint, which regenerates a selected part. It describes one-click LoRA training in Gradio, with 8 songs taking about an hour on an RTX 3090.

Hardware tiers run from under 4 GB of VRAM for basic operation to 24 GB or more for the largest model, with CUDA, ROCm, Intel XPU, Apple Silicon (MLX) and CPU paths. Speed is stated as under 2 seconds per full song on an A100 and under 10 seconds on an RTX 3090. The README positions quality between Suno v4.5 and Suno v5. That is the project's own claim, so listen before you rely on it.

What does Sume's Music Router give you?

POST /v1/music-router/generate takes a prompt of up to 5,000 characters and an optional image_url. Omit model or send sume/music-auto and Sume picks the engine, which is Lyria 3.5 today. Explicit ids from the catalog are lyria-3.5 and lyria-3-pro. Every model charges the same fixed Music price per audio generation, $0.125 at the catalog price, regardless of prompt length.

There is no duration field, no negative prompt, no stems and no repaint. You steer length and structure inside the prompt, for example with section markers such as [0:00-0:30] Intro:. Google's docs say Lyria 3.5 is single-turn: you cannot edit a result with a follow-up prompt, and identical prompts give different results.

ACE-Step 1.5 versus Sume Music Router, read 2026-10-02
FeatureACE-Step 1.5Sume Music Router
Length10 seconds to 10 minutes, set by youSteered in the prompt; no duration field
EditingRepaint, covers, stemsNew take only; no iterative editing
CustomizationLoRA training on your own songsPrompt only; no fine-tuning
CostYour hardware$0.125 per generation, whichever engine
OutputFiles on your machineSume-hosted audio artifact (typically audio/mpeg)

What does a Sume request look like?

Put the structure in the prompt, close with the exclusion, and send an idempotency key so a retry never double-charges.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: score-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "[0:00-0:15] Intro: warm lo-fi hip hop, 84 BPM, C minor, dusty Rhodes. [0:15-0:45] Groove: brushed boom-bap drums, muted trumpet answer. A 45-second track. Instrumental, no vocals."
  }'

Read the audio artifact from the result:

curl https://api.sume.com/v1/jobs/$JOB_ID/status \
  -H "Authorization: Bearer $SUME_API_KEY"

curl https://api.sume.com/v1/jobs/$JOB_ID/result \
  -H "Authorization: Bearer $SUME_API_KEY"

Which should you use?

Use ACE-Step for experiments that need control: repainting a bridge, building a LoRA on your own catalog, or generating hundreds of drafts without per-track billing. You carry the setup, the GPU time and the rights review of whatever you train on.

Use Sume when the music is one deliverable among others, such as a bed under a video. You can drop the audio into a Timeline 1.0 soundtrack with gain_db, loop, fade_out_seconds and duck_db instead of editing it by hand. Either way, check the rights terms for commercial use and the labeling rules for AI-made music before publishing.

A practical way to decide is to run one prompt on both. Write a 45-second brief with tempo, key and an arc, generate three takes locally and three through the router, and compare them on the same speakers. Count the minutes you spent on setup as well as the minutes of audio you got. If the local takes are better and you will make many of them, the setup pays back. If you need one clean take per video, the hosted price is hard to beat.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume