MiniMax-Music3 weights: 5-minute songs, 8 GB, vs Sume Music Router
MiniMax-Music3 is open-weights: 5-minute 32 kHz songs, 8 GB with streaming, a $20M licence rule. What it needs, and what Sume's Music Router does instead.

MiniMax-Music3 is an open-weights song model that turns lyrics and a style description into up to five minutes of 32 kHz stereo audio, and it runs in 8 GB of VRAM only with streaming; a 24 GB card is the comfortable size. If you have no CUDA GPU, or you want a track from one prompt, Sume's Music Router returns one at a fixed $0.125 per generation, with the limits listed below.
Model facts are from the MiniMax-Music3 card and its licence file, read 2026-10-03. Sume facts are from the Music Router docs. The card lists it among Hugging Face's trending text-to-audio models.
What does MiniMax-Music3 need to run?
The card describes an 8B global language model plus a 0.6B local model for structure and acoustic detail. Inputs are lyrics with optional section tags such as [Verse] and [Chorus], and a music description covering genre, BPM, key, vocal details and arrangement. The prompt limit is 5,000 tokens and generation is capped at 9,000 acoustic frames. It needs CUDA.
Memory is the practical constraint, and the card gives three figures.
The card does not list a release date or the languages the vocals support, so test your own language and lyric style on a short section before you generate a full five-minute song. Streaming at 8 GB is the minimum, not a recommendation: the card gives no speed figure for it.
| Setup | VRAM on the card | Notes |
|---|---|---|
| Full precision | 24 GB or more | Comfortable on one large card |
| With CPU offloading | About 22 GB | Trades speed for memory |
| Streaming | 8 GB minimum | The smallest setup the card lists |
| Output | Up to 5 minutes | 32 kHz, 16-bit stereo |
What does the licence allow?
The card's badge reads as a Creative Commons licence, but the LICENSE file is a custom MiniMax-Music3 Community License. As the file reads, commercial use is allowed with three conditions: display "MiniMax-Music3" prominently in the user interface of a product that uses it; get written authorization from MiniMax if yearly revenue from products using it exceeds $20 million; and keep reasonable technical and organizational safeguards against misuse. An Acceptable Use Policy lists prohibited uses such as false information and deepfakes.
That makes it more usable for a small studio than a non-commercial music model, but the UI-display rule is a real obligation, and a track you deliver to a client may need it handled in the contract. Read the file itself; this post is a summary, not advice. For a non-commercial comparison, see MusicGen's weights.
What does Sume's Music Router do instead?
Sume's Music Router takes a prompt of 1 to 5,000 characters and picks the engine when model is omitted or set to sume/music-auto, which is Lyria 3.5 today. Explicit ids are lyria-3.5 and lyria-3-pro. Every generation charges the fixed music price, $0.125, whatever the prompt length (Music 1.0).
It differs from MiniMax-Music3 in ways that matter for a plan. There is no separate lyrics field in the request: structure goes into the prompt, such as a time-stamped section outline. duration and duration_seconds are rejected, so length is steered in words. A non-empty negative_prompt is unsupported. The result can carry result.lyrics, which is model-reported metadata, and job.request.routed_model names the engine that ran.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: music-router-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Warm lo-fi hip hop, 84 BPM, C minor. [0:00-0:30] Intro, Rhodes chords and brushed drums. Instrumental, no vocals. A 2-minute track."
}'Which should you choose?
Choose by hardware, control and obligations.
Neither is a free pass on rights: whatever generates the song, you still need to check what the output may be used for in your channel.
- You own a 24 GB CUDA card, want tagged lyrics and up to five minutes, and can show the required branding: MiniMax-Music3 is worth a test.
- You want one prompt, one job and no GPU: use the Music Router and poll the job with the usual status and result calls from Jobs and results.
- You need a fixed length: say it in the prompt for Sume, or cut afterwards with timeline audio split ranges.
- You need a negative prompt: Sume does not support one on music, so write exclusions positively.
Sources
Related posts
More in Models
- Nano Banana Pro interleaved text and images vs Sume image output
Google documents Nano Banana Pro returning text blocks with illustrations in one answer. Sume's Image API documents image results only; here is the workaround.
- OmniVoice: 600+ languages, CC-BY-NC weights, hosted TTS instead
OmniVoice covers 600+ languages in a 0.6B model, but its weights are CC-BY-NC. What the card says, what it omits, and where a hosted TTS job fits.
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
- Polish, Dutch, Swedish, Turkish text to speech API: Sume pl nl sv tr
Sume's Voices library tags voices pl, nl, sv and tr alongside 12 other languages. What to send for each, and how Eleven v4's list compares.
Written by Sume