Silent reference clip: stt_skipped_silent, no-audio and caption errors
A silent clip does not fail a Sume reference ingest, and STT settles to zero. Video inspect and captions do raise errors on silence.

A silent clip does not fail a Sume reference ingest; the speech-to-text reservation settles to zero and the response carries a warning such as stt_skipped_silent. The two other video tools behave differently: video inspect with transcribe: true returns inspect_source_has_no_audio for a clip without an audio track, and video captions return caption_no_speech when there is nothing audible to caption.
These behaviours come from the Sume docs for Reference ingest, Video inspect and Video captions. Reference ingest is listed as dest first, so confirm it is enabled for your workspace before building on it.
Reference ingest: the quiet path
Reference ingest reads one clip of up to 300 seconds you already hold on media.sume.com. speech.allow_billed_stt reserves the STT per-minute rate but transcribes only when the track is not silent and voice activity detection finds speech; otherwise the charge settles to zero. It defaults to true for purpose: reference_remix and false for the other purposes. The audio block marks a clip silent when loudness is at or below minus 60 LUFS, and the docs say to plan new music and discard the source audio when audio.silent is true.
| Tool | Silent or no-audio clip | What you see |
|---|---|---|
| Reference ingest | Succeeds; STT settles to zero | Warnings stt_skipped_silent, stt_skipped_no_speech or stt_skipped_no_audio_track |
| Video inspect with transcribe | Fails for no audio track | inspect_source_has_no_audio |
| Video captions without your own text | Fails | caption_no_speech, next_action use_overlay_captions |
Video inspect and captions: check first
Video inspect tells you to examine probe.has_audio first, and says a frames: false inspect is enough for that. If it is false, drop transcribe from the call. Video captions work from script_text or speech-to-text, and speech needs audible content. If a clip is silent, burn authored overlay text by sending cues or segments with text, start and end; that path uses no speech-to-text.
A simple decision flow
Before sending a reference clip to any speech step, decide from the clip rather than from the file name:
- No audio track: skip STT, plan new music or voice-over.
- Audio track but silent: the same, and expect
audio.silent: true. - Music only: ingest reports speech absent, so no STT charge, and beats from the music track.
- Speech present: allow STT and give
duration_secondsso the reservation matches the clip.
What it costs when speech is present
STT is $0.01 per audio minute. A 45 second clip with a duration hint of 45 reserves one ceiling minute, so one cent. If you omit the hint, the docs say the reservation is one minute as well. The ingest also bills its media compute, which the docs say is the container seconds times the Modal list price times 1.25 plus the platform fee, never more than the hold, so I give no dollar total for that part.
Sources
Related posts
More in Developers
- Retry or not: a decision table for failed Sume submit calls
Which Sume errors deserve a retry with the same Idempotency-Key, which mean poll the job, and which mean fix the request. One table and a 15-line classifier.
- Revoked a Sume API key and it still works? Up to 15 seconds
After you revoke a Sume API key, the API can keep accepting it for up to 15 seconds by default, because the key lookup is cached. What that means for a leak.
- Ruby verifier for Sume avatar job webhooks (HMAC SHA 256)
A Ruby method that verifies Sume's x-sume-webhook-signature: HMAC SHA 256 over timestamp.body, any sume-v1 entry, five-minute window, empty secret refused.
- Slow down a TTS voiceover: speed 0.6 to 1.5, volume, emotion
Sume TTS accepts generation_config with speed 0.6 to 1.5, volume 0.5 to 2 and a free-text emotion guide. How to retime a voiceover to a video.
Written by Sume