FFmpeg drawtext on a hosted API: why Sume filters refuse it, and cues

Sume video-filter refuses drawtext and subtitles because they read files. To burn text onto a clip, send cues to video-captions: text, start and end seconds.

4 min readSume
All posts

Sume's video filter does not accept drawtext: it is not on the allowlist, and the docs say drawtext, subtitles, movie and lut3d are out, along with anything that reads a file or socket. To put your own text on a clip, send cues to video captions: each cue is text, start and end in seconds, and Sume burns exactly that copy at those times with no speech-to-text.

FFmpeg's description of drawtext is from its filters documentation, read 2026-10-02.

Why is drawtext refused?

FFmpeg's entry says drawtext draws a text string, or text from a specified file, using libfreetype, and takes a font setup. The Sume filter list is limited to filters with only inline options so that nothing on it can read a path. An off-list name returns invalid_filtergraph naming the allowed filters. A free POST /v1/video-filter/check shows that before you pay.

How do cues work?

cues (alias segments) holds 1 to 200 items; each has text up to 400 characters, and start and end from 0 to 60 seconds. Text may contain a newline for a two-line card. cues cannot be combined with words, segments or script_text. Pick a style (slam, punch, tiktok-green, or a Hangul style) and tune it with design. A standalone caption job is $0.20 for videos up to 60 seconds.

The video_url must be a fetchable public HTTPS URL. Poll GET /v1/jobs/:id/status and read the result, or GET /v1/video-captions/:id.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cue-text-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "punch",
    "cues": [{ "text": "Free shipping today", "start": 0.5, "end": 3 }]
  }'

When is cues the wrong tool?

A silent clip with no cues fails as caption_no_speech; the fix is the cue form above.

What each surface takes, from the FFmpeg drawtext entry and the Sume docs, read 2026-10-02.
TaskWhere it goes on SumeNote
Free text on a clipcues on video captions0 to 60 s per cue
Speech to captionsOmit cues; STT runsNeeds audible speech
A logo or image plateTimeline compose overlayOne still on one video
Arbitrary drawtext optionsNot offeredNot on the allowlist

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume