Music API waveform data: get peaks from the audio file yourself

ElevenLabs added optional waveform data to music generation. Sume's music result has no waveform field, so compute peaks from the audio artifact with ffmpeg.

5 min readSume
All posts

A music API waveform is a list of peak values you can draw as bars in a player. ElevenLabs' changelog lists optional waveform visual data for music generation. Sume's music result carries no waveform field, so you compute the peaks yourself from the audio artifact: download the file from the job result, decode it to mono samples with ffmpeg, and take the maximum in each bin.

The ElevenLabs entry is from its API changelog, read 2026-09-29. The Sume result shape is from Music 1.0 and the Music Router docs.

What does a Sume music job return?

A completed job returns Sume-hosted audio artifacts. Each artifact has these fields, and the docs show no waveform among them.

From the Sume Music 1.0 artifact example, read 2026-09-29.
FieldValue in the docs' example
idartifact_...
typeaudio
urlhttps://media.sume.com/artifacts/...
content_typeaudio/mpeg

How do I fetch the audio file?

Poll the job, then read result.artifacts[] for the entry whose type is audio. Its url is a Sume media URL you can download directly.

curl https://api.sume.com/v1/jobs/job_123/result \
  -H "Authorization: Bearer $SUME_API_KEY"

curl -o track.mp3 "https://media.sume.com/artifacts/..."

How do I turn the file into peaks?

This script asks ffmpeg for mono, 8 kHz, 16-bit samples, splits them into 800 bins and prints each bin's peak from 0 to 1. Pass the file path as the argument. Draw the numbers as bar heights.

import array
import json
import subprocess
import sys


def peaks(path, bins=800):
    raw = subprocess.run(
        ["ffmpeg", "-v", "error", "-i", path, "-ac", "1", "-ar", "8000", "-f", "s16le", "-"],
        capture_output=True,
        check=True,
    ).stdout
    samples = array.array("h")
    samples.frombytes(raw)
    step = max(1, len(samples) // bins)
    return [
        round(max(abs(x) for x in samples[i : i + step]) / 32768, 3)
        for i in range(0, len(samples), step)
    ]


print(json.dumps(peaks(sys.argv[1])))

Where should I compute peaks?

Compute them once, when the job completes, and store the array beside the artifact id. A webhook-mode job can trigger that step, and the player then reads a small JSON file instead of decoding audio in the browser. Sume's music docs promise no timing or loudness data; the lyrics field is model-reported metadata, not an audio measurement, so do not use it to draw a waveform.

What does the ElevenLabs changelog say?

Two entries matter here. These are ElevenLabs' own notes, and we did not call its API.

From ElevenLabs' API changelog, read 2026-09-29.
DateEntry
2026-08-31Optional waveform visual data for music generation
2026-09-14Music v2.5 API support

Why decode to mono at 8 kHz?

A player bar needs a few hundred values, not full-quality audio. Mono at 8 kHz keeps the peaks' shape while keeping the decode fast and the array small. Raise the sample rate or the bin count if your player is very wide or you zoom in. The numbers in the script are choices of ours, not a Sume or ElevenLabs specification.

Because the audio artifact is audio/mpeg in the docs' example, ffmpeg does the decoding; no Sume route returns samples.

What should I store next to the peaks?

Keep the peaks array in the same record as the job id and the artifact id, so a player can draw the bars before the audio has downloaded.

  • The job id, so you can fetch the result again.
  • The artifact id and its media.sume.com URL.
  • The bin count and sample rate you used, so a later change does not mix two shapes.
  • The peaks array itself, as JSON.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume