audio_spine_low_fidelity: why a 16 kHz transcript wav warns
Sume warns audio_spine_low_fidelity when a render's audio.url is a 16 kHz mono speech-to-text wav. Detach the track at its source rate instead.

audio_spine_low_fidelity is a warning Timeline 1.0 raises when the render's audio.url is a speech-to-text file, a 16 kHz mono wav, or any spine under 32 kHz, or mono under a stereo video source. Detach the track with audio_detach at the source sample rate and channels, and use that as the spine.
Where the file comes from
Video inspect returns a transcript whose audio_url is a 16 kHz mono pcm_s16le wav. The API schema calls it speech-to-text input, reusable as an STT audio_url or a voice-clone sample, and says it is not a timeline spine, because everything above 8 kHz and the stereo image are gone.
| File | Good for | Not for |
|---|---|---|
| Video inspect transcript audio_url (16 kHz mono) | STT input, voice-clone sample | Render spine |
| Audio detach wav at source rate and channels | Render spine | Larger than STT needs |
| Audio detach 16000 Hz mono | STT input | Render spine |
| A TTS master | Render spine | Not applicable |
The fix
Call POST /v1/audio-detach with format: "wav", leave channels and sample_rate unset so they inherit the source, and pass the returned artifact URL as audio.url on the render. It costs $0.01 per job. The master inherits the spine's rate and channels, so a poor spine produces a poor render.
It is a warning, not a failure
The render still runs, and the master can only carry what the thin spine kept. Check the job's warnings before you ship. The source channels post explains the inherit behaviour in detail.
Sources
Related posts
More in Media tools
- Azure batch transcription wants WAV PCM or FLAC: Sume detach's default
Azure advises lossless WAV (PCM) or FLAC for best transcription quality. Sume audio detach defaults to sample-exact pcm_s16le wav, so the default fits.
- Azure fast transcription: 500 MB, under 5 h, vs Sume's 900 s detach
Azure fast transcription takes audio under 500 MB and under 5 hours. Sume audio detach outputs at most 900 s, so longer audio needs ranges.
- Batch trim clips from a spreadsheet of start and end times
Read start and end columns from a CSV and trim a long video into Shorts: one Video trim call per row at $0.02, with idempotency keys.
- Black bars on a YouTube Short: YouTube says no, so fill the frame
YouTube says uploads should never include letterbox or pillarbox bars, and Shorts take square or vertical files. Reframe 16:9 clips with Sume Timeline fit.
Written by Sume