Silent reference clip: reference_ingest's -60 LUFS gate and STT
reference_ingest marks a clip audio.silent at -60 LUFS or below and skips STT, so a silent reference costs no transcript. What to do next: plan new music.

A reference clip is silent to reference_ingest when its integrated loudness is at or below -60 LUFS, or its true peak is minus infinity. In that case audio.silent is true, the speech-to-text charge settles to zero, and the manifest tells you to plan new music instead of reusing the source audio. Never ask for a transcript of a track that is silent.
The gate, in order
The audio block in the manifest has three parts: a loudness gate, Silero VAD speech presence, and, when the track is music, librosa beats. A transcript is a fourth step that runs only if the earlier ones permit it. When speech.allow_billed_stt is true, Sume reserves the STT 1.0 per-minute rate, then transcribes only if the track is not silent and VAD finds speech. In any other condition the charge settles to zero.
What the manifest says when nothing runs
Skipped steps leave a warning rather than a silent gap. You may see stt_skipped_silent, stt_skipped_no_speech or stt_skipped_no_audio_track. They tell you why there is no transcript, so a missing field is not a failure to chase.
| Case | Warning | STT charge | Next step |
|---|---|---|---|
| Integrated loudness at or below -60 LUFS | stt_skipped_silent | Settles to zero | Plan new music; drop source audio |
| Audible, no speech found | stt_skipped_no_speech | Settles to zero | Beats analysis may apply |
| No audio track | stt_skipped_no_audio_track | Settles to zero | Check with a video_inspect probe |
| Audible with speech | None | Per-minute rate | Words and sentence segments in manifest |
Why the default matters
speech.allow_billed_stt defaults to true when purpose is reference_remix and false for the other purposes, and false opts out. With a silent reference you lose nothing by leaving it on, since the charge settles to zero, but an explicit false makes your intent plain in the request log. The read still pays its own Modal compute, captured at container seconds times the Modal list times 1.25 plus the platform fee, never above the hold.
If you send speech.language_code or duration_seconds on a read that does not transcribe, you get reference_ingest_stt_required. Do not send them on a clip you expect to be silent.
What replaces the audio
A silent source means its audio cannot be carried into the remix. Choose new music for the cut and match it to the manifest's shots, which tile the clip with no gap, rather than to the missing beat. If the clip has music but no speech, the beats in the manifest give you something to match against. Check the loudness by ear on the first pass; a quiet clip that sits just above -60 LUFS is audible and will be read as audio. Record the decision in your brief, for example 'source silent, new score needed', so the person who edits the cut does not search for audio that was never there. A manifest that says silent is evidence, and it is cheap to keep.
Sources
Related posts
More in Media tools
- reference-ingest source_no_video_stream: audio-only file as input
source_no_video_stream means the reference file has no video track, such as an m4a. Send a clip with picture, or use audio detach to work with the sound.
- reference_ingest_stt_required: speech.language_code without STT
speech.language_code and duration_seconds only apply when the read transcribes. Set speech.allow_billed_stt true, use reference_remix, or drop them.
- Reference ingest text_tracks is not caption coverage: 5 sampled frames
text_tracks come from OCR on deduplicated frame states, at most five frames, not every caption. To know if a Short is captioned, check frames yourself.
- Reframe 16:9 to 9:16: crop fractions for left, center, right
A 16:9 to 9:16 crop is a width of 0.3164 of the frame. Use x 0, 0.3418 or 0.6836 for a left, centered or right subject in Sume video filter.
Written by Sume