Silent reference clip: stt_skipped_silent, no-audio and caption errors

A silent clip does not fail a Sume reference ingest, and STT settles to zero. Video inspect and captions do raise errors on silence.

5 min readSume
All posts

A silent clip does not fail a Sume reference ingest; the speech-to-text reservation settles to zero and the response carries a warning such as stt_skipped_silent. The two other video tools behave differently: video inspect with transcribe: true returns inspect_source_has_no_audio for a clip without an audio track, and video captions return caption_no_speech when there is nothing audible to caption.

These behaviours come from the Sume docs for Reference ingest, Video inspect and Video captions. Reference ingest is listed as dest first, so confirm it is enabled for your workspace before building on it.

Reference ingest: the quiet path

Reference ingest reads one clip of up to 300 seconds you already hold on media.sume.com. speech.allow_billed_stt reserves the STT per-minute rate but transcribes only when the track is not silent and voice activity detection finds speech; otherwise the charge settles to zero. It defaults to true for purpose: reference_remix and false for the other purposes. The audio block marks a clip silent when loudness is at or below minus 60 LUFS, and the docs say to plan new music and discard the source audio when audio.silent is true.

What each Sume tool does with silence (per Sume docs, read 2026-10-10)
ToolSilent or no-audio clipWhat you see
Reference ingestSucceeds; STT settles to zeroWarnings stt_skipped_silent, stt_skipped_no_speech or stt_skipped_no_audio_track
Video inspect with transcribeFails for no audio trackinspect_source_has_no_audio
Video captions without your own textFailscaption_no_speech, next_action use_overlay_captions

Video inspect and captions: check first

Video inspect tells you to examine probe.has_audio first, and says a frames: false inspect is enough for that. If it is false, drop transcribe from the call. Video captions work from script_text or speech-to-text, and speech needs audible content. If a clip is silent, burn authored overlay text by sending cues or segments with text, start and end; that path uses no speech-to-text.

A simple decision flow

Before sending a reference clip to any speech step, decide from the clip rather than from the file name:

  • No audio track: skip STT, plan new music or voice-over.
  • Audio track but silent: the same, and expect audio.silent: true.
  • Music only: ingest reports speech absent, so no STT charge, and beats from the music track.
  • Speech present: allow STT and give duration_seconds so the reservation matches the clip.

What it costs when speech is present

STT is $0.01 per audio minute. A 45 second clip with a duration hint of 45 reserves one ceiling minute, so one cent. If you omit the hint, the docs say the reservation is one minute as well. The ingest also bills its media compute, which the docs say is the container seconds times the Modal list price times 1.25 plus the platform fee, never more than the hold, so I give no dollar total for that part.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume