Bannerbear video tools mapped to Sume endpoints, gap by gap
Bannerbear lists eleven video tools. Sume covers trim, crop, concat, captions and color via filters; picture-in-picture and GIF previews are not documented.

Bannerbear's v5 API lists eleven video tools. Sume covers five of them with dedicated endpoints (trim, crop, concat, captions, and colour or pixel filters), and the docs do not describe a picture-in-picture video overlay, skin softening or GIF preview.
Tool names come from the Bannerbear API reference; Sume mappings come from the model docs, and anything I could not find there is marked as not documented.
What are Bannerbear's video tools?
Bannerbear exposes asynchronous POST endpoints under /v5/tools/: trim_video, concat_videos, resize_video, crop_video, overlay_video, overlay_image, subtitle_video, add_audio, apply_color_filter, soften_video and create_gif_preview. Each returns 202, then you poll GET /v5/tool_jobs/:uid, whose status is pending, running, completed or failed, with a progress field from 0 to 100 and an outputs object when done.
Which Sume endpoint does each tool map to?
Sume splits the same work into small prep endpoints and one assembly endpoint. Each takes a workspace media.sume.com URL, so import first.
| Bannerbear tool | Sume surface | Notes |
|---|---|---|
| trim_video | Video trim, $0.02 per job | start plus end or duration |
| concat_videos | Timeline 1.0, $0.10 per ceil output minute | Needs an audio spine or audio.mode: silence; six transition types |
| resize_video | Timeline output.width/height, slot fit | cover, contain, stretch, blur; even sizes 256 to 2160 |
| crop_video | Video filter crop op, $0.02 | Fractions of the frame, not pixels |
| overlay_video | Not documented | Timeline compose takes one still plus one video |
| overlay_image | Timeline compose, $0.02 | stack or overlay; one still, one video; up to 300 s |
| subtitle_video | Video captions, $0.20 up to 60 s | Styles, cues for authored text |
| add_audio | Timeline audio spine and soundtrack | Soundtrack gain, loop, fade and duck |
| apply_color_filter | Video filter filtergraph | Allowlisted filters only; dim op |
| soften_video | Not documented | |
| create_gif_preview | Not documented | Video frames returns stills, not GIFs |
Where does the shape differ?
Two differences matter in code. First, Bannerbear tools take video_url pointing anywhere fetchable; Sume's media prep endpoints take only this workspace's media.sume.com artifacts, and off-host URLs are rejected at admit. Import with POST /v1/media-imports first. Video captions is the exception and accepts a public HTTPS URL.
Second, Sume never accepts ffmpeg options. A vf, filter, codec or crf field returns ffmpeg_fields_rejected, and the filtergraph for video filter is capped at 2048 characters and 32 filters. A status read is GET /v1/jobs/:id/status with queued, processing, completed, failed or canceled, not a tool-specific job route.
Example: the trim call
Where Bannerbear would use POST /v5/tools/trim_video, Sume uses one call with an idempotency key.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: trim-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"start": 2,
"end": 10
}'Should you move?
If your workflow is templated social graphics plus a few video tools, Bannerbear's single surface is simpler. If you already generate video with Sume and want trim, filter and assembly to live next to it, the Sume endpoints save a hop. Neither is cheaper in general, since Bannerbear's page I read does not list video tool prices. Check the plan or catalog before you commit volume.
What gaps should you plan around?
Three of Bannerbear's tools have no documented Sume counterpart: overlay_video (picture-in-picture of a second video), soften_video and create_gif_preview. Timeline compose takes exactly one still and one video, so a video-on-video overlay is not covered, and video frames returns JPEG or PNG stills rather than animation.
A second gap is units. Bannerbear's crop_video takes pixel coordinates, while Sume's crop op takes fractions of the frame between 0 and 1, with width and height from 0.05, so convert before you port. Timeline sizes must be even integers from 256 to 2160, and an odd source height is a common cause of failures; see the ffmpeg odd height guide.
- Convert crop pixels to fractions.
- Import every source to
media.sume.comfirst. - Video captions accepts a public HTTPS URL.
Sources
Related posts
More in Comparisons
- Beatoven maestro music and SFX API vs Sume Music Router
Beatoven's API makes music and sound effects from text. Sume's Music Router makes one track per call at a fixed $0.125 and has no SFX route. The differences.
- Best TTS model right now: the leaderboard versus Sume's router
Eleven v4 leads the Artificial Analysis TTS board today. What that means if you generate speech through Sume, whose router serves Sonic models only.
- Cartesia acceptable use: the explicit consent clause for clones
Cartesia's policy allows only your own voice or others' with explicit consent. What that means for a voice you clone on Sume, which runs on Sonic.
- Cartesia Ink STT hours per plan vs Sume STT $0.60 per hour
Cartesia plans include about 9 to 741 hours of Ink speech-to-text. That works out near $0.40 to $0.54 an hour; Sume STT is $0.60 an hour. Concurrency compared.
Written by Sume