ACE-Step 1.5 vocal-to-background music: what Sume does instead

ACE-Step 1.5 lists a vocal-to-background-music mode, covers and repainting under MIT. Sume's music jobs are prompt-only, so here is how to get a vocal-free bed.

4 min readSume
All posts

If you want to turn a song with vocals into a background track, ACE-Step 1.5 lists that as a feature, and Sume does not. The model card names "vocal-to-background music conversion" next to text-to-music, covers, repainting and fine-tuning, under an MIT license. Sume's music jobs take a text prompt and an optional image, so on Sume you ask for a bed without vocals instead of converting an existing song.

What does ACE-Step 1.5 say it can do?

The card is short and specific. These are its claims, not measurements we ran.

ACE-Step 1.5 model card, https://huggingface.co/ACE-Step/Ace-Step1.5, read 2026-10-03.
ItemModel card says
LicenseMIT
MemoryRuns locally with less than 4GB of VRAM
SpeedFull song in under 2 seconds on an A100, under 10 seconds on an RTX 3090
Languages50+
LengthShort loops up to 10-minute compositions
FeaturesText-to-music, covers, repainting and editing, vocal-to-background music, fine-tuning
Training dataLicensed, royalty-free or no-copyright, and synthetic data

What does Sume accept as input?

Music 1.0 takes prompt (1 to 5000 characters) and an optional public HTTPS image_url. There is no audio input, no seed and no duration field, and the Music Router uses the same body plus an optional model. So an existing song cannot be uploaded and converted.

How do you get a vocal-free bed on Sume?

Write the exclusion into the positive prompt. The docs say Lyria does not support negative prompting and that a non-empty negative_prompt returns HTTP 400. Close the brief with "Instrumental, no vocals.", and add "no spoken word" when the bed goes under narration. Then check the result by ear, since the docs treat the brief as direction, not a guaranteed setting.

Tempo, key and instruments help the match: the music page suggests a number for tempo (for example 72 BPM), a key and two to four instruments with texture.

Which should you use?

Use ACE-Step locally when the job is converting or repainting a track you already have the rights to and you have a GPU. Use a Sume Music Router job when you need a new bed with a Sume-hosted URL, billed at the fixed Music price, that can go into a Timeline soundtrack with loop, gain_db and duck_db. Check the rights to any song you feed a converter, whatever the converter's license says.

What does the speed claim change?

The card says a full song in under 2 seconds on an A100 and under 10 seconds on an RTX 3090. That matters when you want dozens of takes to audition. It does not tell you the takes are good, and the numbers are the model card's, not ours. On Sume, a music job is asynchronous: the docs say sync and subscribe modes wait up to 30 seconds, and a song may take longer, so poll the job instead of waiting on the request.

Either way, plan to listen. A bed that sounds right alone can fight a voice, which is what the duck_db field in the Timeline soundtrack is for.

What does the training-data line tell you?

The card says the training data comprised licensed data, royalty-free or no-copyright data and synthetic data, and that generated music is available for commercial purposes. That is the vendor's statement, and it is worth saving with the date you read it. It is not a guarantee about any single output, and it says nothing about the songs you might feed into a conversion mode.

Sources

Related posts

More in Models

All Models posts

Written by Sume