FFmpeg whisper filter vs a hosted transcript call for video
FFmpeg's whisper filter needs whisper.cpp and a model file you manage. Sume's video-inspect transcribe returns words and sentence segments for $0.01 a minute.

FFmpeg's whisper filter runs speech recognition inside an ffmpeg command, but only if your build was configured with --enable-whisper against the whisper.cpp library and you point it at a downloaded model file. Sume's video inspect does the same job as a hosted call: send transcribe: true for a clip already on media.sume.com and read back text, word timings and optional sentence segments, at a public rate of $0.01 per audio minute.
The FFmpeg side comes from the FFmpeg filters documentation and the FFmpeg release news, both read on 2026-10-02. The news page lists the Whisper filter among the filters added in the 8.0 release (22 August 2025). Sume's side comes from its video inspect, video captions and audio detach docs. This is not a quality benchmark: nothing here claims one engine transcribes better than the other.
What does the FFmpeg whisper filter need before it runs?
The filter documentation says it needs the whisper.cpp library as a prerequisite, enabled with ./configure --enable-whisper. The model option, the file path of a downloaded whisper.cpp model, is mandatory. Everything else has a default: language defaults to auto, queue to 3 seconds, use_gpu to true and format to text.
That makes the first hour of work about packaging rather than transcription: a build that includes the filter, a model file on disk, and a decision about GPU use. If your editing automation already shells out to ffmpeg on a machine you control, that is a reasonable trade. If it runs in a serverless function or a CI job, you now own a model download and a build flag.
destinationwrites output to a file or any FFmpeg AVIO URL; without it the text is only logged as info messages and set in thelavfi.whisper.textframe metadata.formatcan betext,srtorjson.max_lensplits segments by word to keep subtitle lines short.vad_modelloads a Silero voice-activity model and the docs suggest raisingqueue(for example to 20) when you use it.
How does the queue setting change what you get?
The docs are direct about the trade-off: a small queue processes the audio more often, so transcription quality is lower and the processing cost is higher; a large value such as 10 to 20 seconds is more accurate and uses less CPU, but adds latency, so it is not suited to real-time streams. The default of 3 seconds is tuned toward the live case, which is not what most editing automation does.
Language handling has a similar catch. auto is currently an alias for eval, which re-runs language detection on every queued chunk at the cost of an extra encoder pass and may switch language mid-file. The docs say to use eval or lock explicitly if the behaviour matters. Sume's equivalent is a single optional language_code hint on the inspect call; omit it for auto-detection.
What does the hosted equivalent look like?
Video inspect reads one clip already owned by the workspace, so import it first with POST /v1/media-imports. Probe and stills are unbilled; only the transcript reserves, at $0.01 per audio minute, and omitting duration_seconds reserves one minute with a maximum hint of 600 seconds. Add segmentation.mode: "sentence" to also get gapless sentence segments shaped like caption lines.
A silent clip fails with inspect_source_has_no_audio, so check probe.has_audio first with a frames: false inspect. Sume does not return an SRT file from this call; you get structured words[] and segments[] and decide what to do with them.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: inspect-transcript-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"language_code": "en",
"segmentation": { "mode": "sentence" }
}'Which one fits which job?
Choose FFmpeg when you already run ffmpeg on your own hardware and want an SRT on disk. Choose the hosted call when the pipeline is an agent or a webhook-driven job that should not carry a model file.
| Question | FFmpeg whisper filter | Sume video inspect |
|---|---|---|
| Setup | Build with --enable-whisper, install whisper.cpp, download a model | Import the clip, send transcribe: true |
| Output | text, srt or json to a file or URL | words[], optional sentence segments[], audio_url |
| Language | language option; auto aliases eval | Optional language_code hint; omit to auto-detect |
| Cost | Your own compute | $0.01 per audio minute (confirm in /v1/catalog) |
| Length | Whatever your machine handles | Source up to 1800 s; billed hint up to 600 s |
| Burn captions in | Separate subtitles filter pass you configure | Pass the words to video captions as words or cues |
What does Sume not do here?
Sume does not expose the ffmpeg command line. Fields such as vf, filter, ffmpeg or cmd are rejected with ffmpeg_fields_rejected, and the server compiles the work itself. Sume also has no whisper-style streaming transcription: inspect is a request, a job and a result. To get words onto the picture, hand them to video captions as words, or as cues for phrase-level cards, which skips speech-to-text on that call; the caption job costs $0.20 for videos up to 60 seconds under the current estimate. For the cost side of hosted STT alone, see Whisper API cost per minute versus Sume STT.
Sources
Related posts
More in Comparisons
- FFmpeg xfade has 59 transitions: which does Sume Timeline take?
FFmpeg's xfade page lists 59 transition values, including custom. Sume Timeline 1.0 accepts six: fade, wipeleft, wiperight, slideup, slidedown and dissolve.
- FLUX 3 Image reference size: 256 px to 16 MP, and Sume URL rules
BFL says each FLUX 3 Image reference is 256x256 px to 16 MP, 1 to 10 images. Sume lists FLUX.2 ids; its reference rules are public HTTPS and a catalog count.
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra's developer platform: API, SDK, CLI and MCP vs Sume
Hedra opened its models through an API, SDKs, a CLI and MCP on August 4, 2026. A map of what each surface covers and what Sume offers for avatar work.
Written by Sume