Scribe accepts MP4 and MOV; Sume STT needs an audio URL: detach first

ElevenLabs Scribe takes video files such as MP4 and MOV directly. Sume STT wants a public HTTPS audio_url, so a video goes through audio detach first at $0.01.

5 min readSume
All posts

ElevenLabs Scribe accepts video files directly, including MP4, MOV, MKV and WebM, while Sume's STT route takes only a public HTTPS audio_url. For a video you hold on Sume, the shortest correct path is audio detach to a 16 kHz mono WAV ($0.01 per job), then an STT job on the resulting audio, or transcribe: true on a video inspect if you also want stills.

The Scribe format list is from ElevenLabs' speech-to-text page (read 2026-10-10). Sume's request and the detach settings are from the Sume OpenAPI contract and Audio detach.

The formats ElevenLabs lists

The page lists audio formats AAC, AIFF, OGG, MP3, OPUS, WAV, FLAC, M4A and WebM, and video formats MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG and 3GPP, with a 3 GB ceiling and a 10-hour maximum for standard mode. Sume's STT schema does not list formats at all; it says to provide a public HTTPS audio URL and prefer a Sume media or attachment URL. Because the list is silent, do not assume Sume STT reads an MP4 URL directly.

Inputs compared (ElevenLabs read 2026-10-10; Sume per OpenAPI and docs.sume.com)
InputElevenLabs ScribeSume STT 1.0
Audio fileMany formats listedPublic HTTPS audio_url; formats not listed in the contract
Video fileMP4, MOV, MKV and moreDetach the audio first, or use video inspect with transcribe: true
Size or length3 GB and 10 hoursduration_seconds 1 to 600 per job
Preparation costNoneDetach $0.01 per job

The detach step

Audio detach takes one workspace media.sume.com video, so import the file first with a media import. Set format: "wav", channels: "mono" and sample_rate: 16000; the page calls that the STT shape. The source may be up to 1,800 seconds, and the output up to 900 seconds, so use a range for longer videos. A source with no audio track fails with detach_source_has_no_audio, and the docs suggest checking probe.has_audio with a video inspect first.

The result gives you a new audio_url on media.sume.com. Pass it to the STT job along with language_code if you know the language.

Or let video inspect do both

Video inspect with transcribe: true runs Sume STT 1.0 on the audio of a clip and bills the same $0.01 per audio minute. It reserves one minute when duration_seconds is absent, and the hint maxes at 600 seconds. If you only need a transcript and stills, that is one call instead of two. If you need the audio file as a reusable artifact, such as for a re-voice, detach is better because you keep the WAV.

For a 4-minute clip, the detach route costs $0.01 for detach plus $0.04 of STT, or $0.05. The inspect route costs the STT minutes plus its own media compute, which the docs bill separately by Modal compute, so it is not directly comparable.

A checklist before you submit

A few things trip people on the first run:

  • Is the URL public HTTPS, or a Sume-hosted media URL?
  • Did you send duration_seconds? Without it the job reserves one minute.
  • Is the clip over 600 seconds? Split it into windows.
  • Does the video have audio at all? Check probe.has_audio.
  • Do you need words? They always come back; sentence segments need segmentation.mode: sentence.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume