Silent reference clip: reference_ingest's -60 LUFS gate and STT

reference_ingest marks a clip audio.silent at -60 LUFS or below and skips STT, so a silent reference costs no transcript. What to do next: plan new music.

5 min readSume
All posts

A reference clip is silent to reference_ingest when its integrated loudness is at or below -60 LUFS, or its true peak is minus infinity. In that case audio.silent is true, the speech-to-text charge settles to zero, and the manifest tells you to plan new music instead of reusing the source audio. Never ask for a transcript of a track that is silent.

The gate, in order

The audio block in the manifest has three parts: a loudness gate, Silero VAD speech presence, and, when the track is music, librosa beats. A transcript is a fourth step that runs only if the earlier ones permit it. When speech.allow_billed_stt is true, Sume reserves the STT 1.0 per-minute rate, then transcribes only if the track is not silent and VAD finds speech. In any other condition the charge settles to zero.

What the manifest says when nothing runs

Skipped steps leave a warning rather than a silent gap. You may see stt_skipped_silent, stt_skipped_no_speech or stt_skipped_no_audio_track. They tell you why there is no transcript, so a missing field is not a failure to chase.

reference_ingest audio outcomes, from Sume docs read 2026-10-05
CaseWarningSTT chargeNext step
Integrated loudness at or below -60 LUFSstt_skipped_silentSettles to zeroPlan new music; drop source audio
Audible, no speech foundstt_skipped_no_speechSettles to zeroBeats analysis may apply
No audio trackstt_skipped_no_audio_trackSettles to zeroCheck with a video_inspect probe
Audible with speechNonePer-minute rateWords and sentence segments in manifest

Why the default matters

speech.allow_billed_stt defaults to true when purpose is reference_remix and false for the other purposes, and false opts out. With a silent reference you lose nothing by leaving it on, since the charge settles to zero, but an explicit false makes your intent plain in the request log. The read still pays its own Modal compute, captured at container seconds times the Modal list times 1.25 plus the platform fee, never above the hold.

If you send speech.language_code or duration_seconds on a read that does not transcribe, you get reference_ingest_stt_required. Do not send them on a clip you expect to be silent.

What replaces the audio

A silent source means its audio cannot be carried into the remix. Choose new music for the cut and match it to the manifest's shots, which tile the clip with no gap, rather than to the missing beat. If the clip has music but no speech, the beats in the manifest give you something to match against. Check the loudness by ear on the first pass; a quiet clip that sits just above -60 LUFS is audible and will be read as audio. Record the decision in your brief, for example 'source silent, new score needed', so the person who edits the cut does not search for audio that was never there. A manifest that says silent is evidence, and it is cheap to keep.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume