Blind-test AI music models with one brief: Sume Music Router ids

New music models launch monthly (Suno v6, ElevenLabs Music v2.5). A simple blind test with one brief, and the three routable ids Sume's router accepts.

5 min readSume
All posts

To choose between AI music models, give each the same written brief, strip the labels, and have two people rank the takes. Models are changing fast: TechCrunch reported Suno's v6 family on 2026-09-09, and a changelog roundup reports ElevenLabs Music v2.5 as default from 2026-09-11. Neither source gives a head-to-head score, so a blind test on your own brief is the only evidence that counts for your use.

Sume can be one of the contestants: its Music Router lets you pin one of three ids, and every generation is a fixed-price job.

How do I set up the test?

Write one brief with Sume's seven axes: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, one named moment in the arc, and the era or production style, closed with "Instrumental, no vocals" if you want no singing. Send the same text to every tool, ask for the same length in words, and generate three takes each. The Music docs note that these axes are creative directions, not guaranteed settings, so judge the audio.

Blind-test plan, with Sume Music Router ids from the docs, read 2026-10-02.
StepWhat to doSume detail
1. BriefOne text, seven axesPrompt up to 5000 characters
2. TakesThree per tool, same length wordingLength is stated in the prompt; no duration field
3. EnginesPin explicit ids where offeredsume/music-auto, lyria-3.5, lyria-3-pro
4. BlindRename files with random codesKeep the job id in your key
5. RankTwo listeners rank independentlyStore the scores with the file

Why pin an id instead of using auto?

Because sume/music-auto resolves to whatever Sume picks, which is Lyria 3.5 today and may change. A test only tells you about the engine you tested. Pin lyria-3.5 or lyria-3-pro for a fair test, and read job.request.routed_model on each result to confirm what ran. Each id charges the fixed Music price per generation, so the test costs the same per take whichever you pin.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: blind-test-a-take-1" \
  -d '{"model":"lyria-3-pro",
       "prompt":"Hushed neo-soul nocturne, 72 BPM, D minor. Rhodes, soft sub bass, brushed snare. A 30-second track. Instrumental, no vocals."}'

How do I read the result?

Rank on fit to the picture, not on how impressive the first ten seconds are, because most music in video sits under speech or action. Check the ending too: a track that stops abruptly needs a fade, which a Timeline render can add. Record the winner with the date and model, since the answer can change when either vendor ships a new version.

The Music Router page lists the ids and request fields, and Music 1.0 has the brief template.

What should I do?

Run the test once per quarter, or when a vendor changes its default. Keep the files and scores so the next test has a baseline.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume