Pika Music: lyrics, voice and track references vs Sume Music Router
Pika Music takes prompts, tagged lyrics, a voice reference and a music reference. Sume's Music Router is prompt-first on Lyria 3.5. See what that changes.

Pika Music accepts four kinds of input, alone or combined: a text prompt, lyrics with section tags, a voice reference and a music reference. Sume's Music Router takes a prompt and an optional image, runs Lyria 3.5 by default, and has no lyrics, voice-reference or audio-reference fields, so lyric-led or reference-led songs are the main thing you give up.
Here is what each side publishes, and how to steer Sume's route when you want structure.
What does Pika Music take as input?
Pika's page lists the four modalities. The prompt describes genre, mood, instrumentation, vocal style and arrangement. Lyrics use section tags such as verse and chorus to shape phrasing. A voice reference guides vocal timbre, delivery and energy. A music reference sets instrumentation, texture and overall feel.
Pika reports a 90-second song in 6.21 seconds on average, and the audio-models post says songs can run up to 6 minutes. For voice and music references the page requires the user to acknowledge that they have the rights to whatever they upload. Access is through the Pika API Club and its playground. The quality scores on the page are Pika's own benchmark results.
What does the Sume Music Router accept?
The Music Router request is a prompt of 1 to 5,000 characters, an optional public HTTPS image_url, and an optional model. The default id is sume/music-auto, which resolves to Lyria 3.5 today, with lyria-3.5 and lyria-3-pro available as explicit ids. The usual mode, webhook_url and metadata fields apply.
| Input | Pika Music | Sume Music Router |
|---|---|---|
| Text prompt | Yes | Yes, 1-5,000 characters |
| Tagged lyrics | Yes | No lyrics field; write lyrics into the prompt |
| Voice reference | Yes, with a rights acknowledgement | No |
| Music reference | Yes, with a rights acknowledgement | No |
| Image conditioning | Not listed | Optional image_url |
| Duration field | Not listed | Rejected; steer length in the prompt |
How do you get structure out of a prompt-only route?
Sume's docs recommend naming the length and using section markers inside the prompt, for example a 2-minute track or [0:00-0:30] Intro. Put exclusions in the positive text, because a non-empty negative_prompt is not supported. Say instrumental with no vocals when you want a bed.
Duration and negative_prompt are rejected rather than ignored, so a client that sends them will see a 400 and not a silent change. A minimal call:
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: music-001" \
-d '{"model":"sume/music-auto","prompt":"Warm lo-fi hip hop, 84 BPM. [0:00-0:15] Rhodes and brushed drums. [0:15-0:30] Add muted trumpet. A 30-second instrumental, no vocals."}'Which one fits which job?
Choose Pika Music when the lyrics are fixed, when you need a sung line in a specific timbre, or when a reference track defines the sound. Those are the cases where Sume's prompt-only route cannot be steered closely.
Choose the Sume Music Router when the music is a bed under a Sume video or avatar render and you want the same wallet, job envelope and timeline in one place. Every Music Router model charges the fixed Music price per generation, so the cost of a retry is predictable.
Whichever you use, music for ads has licence and platform rules that sit outside the model. Read the terms for the plan you are on before you publish, and keep the generated file with its job id.
What should you verify on each side first?
On Pika, the page says Music is available through the API Club and its playground, but it lists no price on the Music page itself. The audio-models post says up to 10x cheaper than comparable models, which is a relative claim, so get the number from the plan you are on before you plan a catalogue of tracks.
On Sume, confirm which engine ran. The job echoes the id you requested, and the routed catalog engine appears in the job request as routed_model, so a request sent to sume/music-auto will show lyria-3.5 there today. That field is how you prove which engine made a track if the default changes later.
Finally, read lyrics back from the result. Sume returns lyrics or a section map in the result when the model reports one, and the docs say that is model-reported metadata, not a measurement of the audio, so listen to the file and do not rely on the text alone. For the full request fields see the Music Router page.
Sources
Related posts
More in Models
- Pika Speech: 5-minute requests, 5-second clones vs Sume TTS
Pika Speech lists 5 minutes per request, 48 kHz output and a 5-second clone at $0.01 a minute. What that means next to Sume's TTS Router and its limits.
- Qwen-Image 2.0 Pro is Alibaba's pick: which Qwen ids does Sume list?
Alibaba recommends qwen-image-2.0-pro. Sume's image catalog lists qwen/qwen-image and qwen/qwen-image-max in the repo; check the live catalog for more.
- Qwen-Image negative_prompt 500 characters vs a Sume image request
Alibaba's Qwen-Image API takes a negative_prompt up to 500 characters. Sume's image request table has no such field, so rewrite exclusions into the prompt.
- Reasoning effort levels: GPT-6.1 Sol, Claude, Grok 4.7, DeepSeek
Effort names differ by vendor: five levels on GPT-6.1 Sol and Claude, four on Grok 4.7, three on DeepSeek. What Sume's picker shows and which rows have no knob.
Written by Sume