audio_spine_low_fidelity: why a 16 kHz transcript wav warns

Sume warns audio_spine_low_fidelity when a render's audio.url is a 16 kHz mono speech-to-text wav. Detach the track at its source rate instead.

3 min readSume
All posts

audio_spine_low_fidelity is a warning Timeline 1.0 raises when the render's audio.url is a speech-to-text file, a 16 kHz mono wav, or any spine under 32 kHz, or mono under a stereo video source. Detach the track with audio_detach at the source sample rate and channels, and use that as the spine.

Where the file comes from

Video inspect returns a transcript whose audio_url is a 16 kHz mono pcm_s16le wav. The API schema calls it speech-to-text input, reusable as an STT audio_url or a voice-clone sample, and says it is not a timeline spine, because everything above 8 kHz and the stereo image are gone.

Which audio file suits which job (read 2026-10-03)
FileGood forNot for
Video inspect transcript audio_url (16 kHz mono)STT input, voice-clone sampleRender spine
Audio detach wav at source rate and channelsRender spineLarger than STT needs
Audio detach 16000 Hz monoSTT inputRender spine
A TTS masterRender spineNot applicable

The fix

Call POST /v1/audio-detach with format: "wav", leave channels and sample_rate unset so they inherit the source, and pass the returned artifact URL as audio.url on the render. It costs $0.01 per job. The master inherits the spine's rate and channels, so a poor spine produces a poor render.

It is a warning, not a failure

The render still runs, and the master can only carry what the thin spine kept. Check the job's warnings before you ship. The source channels post explains the inherit behaviour in detail.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume