FFmpeg whisper filter vs a hosted transcript call for video

FFmpeg's whisper filter needs whisper.cpp and a model file you manage. Sume's video-inspect transcribe returns words and sentence segments for $0.01 a minute.

5 min readSume
All posts

FFmpeg's whisper filter runs speech recognition inside an ffmpeg command, but only if your build was configured with --enable-whisper against the whisper.cpp library and you point it at a downloaded model file. Sume's video inspect does the same job as a hosted call: send transcribe: true for a clip already on media.sume.com and read back text, word timings and optional sentence segments, at a public rate of $0.01 per audio minute.

The FFmpeg side comes from the FFmpeg filters documentation and the FFmpeg release news, both read on 2026-10-02. The news page lists the Whisper filter among the filters added in the 8.0 release (22 August 2025). Sume's side comes from its video inspect, video captions and audio detach docs. This is not a quality benchmark: nothing here claims one engine transcribes better than the other.

What does the FFmpeg whisper filter need before it runs?

The filter documentation says it needs the whisper.cpp library as a prerequisite, enabled with ./configure --enable-whisper. The model option, the file path of a downloaded whisper.cpp model, is mandatory. Everything else has a default: language defaults to auto, queue to 3 seconds, use_gpu to true and format to text.

That makes the first hour of work about packaging rather than transcription: a build that includes the filter, a model file on disk, and a decision about GPU use. If your editing automation already shells out to ffmpeg on a machine you control, that is a reasonable trade. If it runs in a serverless function or a CI job, you now own a model download and a build flag.

  • destination writes output to a file or any FFmpeg AVIO URL; without it the text is only logged as info messages and set in the lavfi.whisper.text frame metadata.
  • format can be text, srt or json.
  • max_len splits segments by word to keep subtitle lines short.
  • vad_model loads a Silero voice-activity model and the docs suggest raising queue (for example to 20) when you use it.

How does the queue setting change what you get?

The docs are direct about the trade-off: a small queue processes the audio more often, so transcription quality is lower and the processing cost is higher; a large value such as 10 to 20 seconds is more accurate and uses less CPU, but adds latency, so it is not suited to real-time streams. The default of 3 seconds is tuned toward the live case, which is not what most editing automation does.

Language handling has a similar catch. auto is currently an alias for eval, which re-runs language detection on every queued chunk at the cost of an extra encoder pass and may switch language mid-file. The docs say to use eval or lock explicitly if the behaviour matters. Sume's equivalent is a single optional language_code hint on the inspect call; omit it for auto-detection.

What does the hosted equivalent look like?

Video inspect reads one clip already owned by the workspace, so import it first with POST /v1/media-imports. Probe and stills are unbilled; only the transcript reserves, at $0.01 per audio minute, and omitting duration_seconds reserves one minute with a maximum hint of 600 seconds. Add segmentation.mode: "sentence" to also get gapless sentence segments shaped like caption lines.

A silent clip fails with inspect_source_has_no_audio, so check probe.has_audio first with a frames: false inspect. Sume does not return an SRT file from this call; you get structured words[] and segments[] and decide what to do with them.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: inspect-transcript-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false,
    "transcribe": true,
    "language_code": "en",
    "segmentation": { "mode": "sentence" }
  }'

Which one fits which job?

Choose FFmpeg when you already run ffmpeg on your own hardware and want an SRT on disk. Choose the hosted call when the pipeline is an agent or a webhook-driven job that should not carry a model file.

FFmpeg whisper filter vs Sume video inspect transcribe, from each vendor's own docs, read 2026-10-02
QuestionFFmpeg whisper filterSume video inspect
SetupBuild with --enable-whisper, install whisper.cpp, download a modelImport the clip, send transcribe: true
Outputtext, srt or json to a file or URLwords[], optional sentence segments[], audio_url
Languagelanguage option; auto aliases evalOptional language_code hint; omit to auto-detect
CostYour own compute$0.01 per audio minute (confirm in /v1/catalog)
LengthWhatever your machine handlesSource up to 1800 s; billed hint up to 600 s
Burn captions inSeparate subtitles filter pass you configurePass the words to video captions as words or cues

What does Sume not do here?

Sume does not expose the ffmpeg command line. Fields such as vf, filter, ffmpeg or cmd are rejected with ffmpeg_fields_rejected, and the server compiles the work itself. Sume also has no whisper-style streaming transcription: inspect is a request, a job and a result. To get words onto the picture, hand them to video captions as words, or as cues for phrase-level cards, which skips speech-to-text on that call; the caption job costs $0.20 for videos up to 60 seconds under the current estimate. For the cost side of hosted STT alone, see Whisper API cost per minute versus Sume STT.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume