Google's GenMedia video editor skill vs Sume trim, filter and Timeline
Google's open GenMedia video editor skill pairs Veo 3.1 with ffmpeg steps. Which of its steps map to Sume's trim, filter and Timeline jobs, and which do not.

Google publishes an open agent skill called genmedia-video-editor that generates cinematic video with Veo 3.1 and then edits it with ffmpeg steps. Sume's docs describe no single skill like it, but its separate trim, filter and Timeline jobs cover some of the same ground, and its hosted MCP server exposes them to an agent.
This page reads the skill's SKILL.md and the Veo 3.1 Lite launch post on Google Cloud, both read on 2026-10-02, plus Sume's Timeline, Video trim and Video filter docs.
What does Google's skill do?
The SKILL.md describes video editing and composition for cinematic video. Its capabilities are generation with Veo 3.1 models, image overlay on video such as logos and watermarks, video to GIF conversion, multi-clip concatenation, and audio and video synchronization. The generation tools it names are text to video, image to video and video extension.
It also gives rules. Clips must have matching dimensions and frame rates before they are merged, you should check media information before a complex operation, and a good prompt follows cinematography, subject, action, context and style. GIF conversion runs in two passes and defaults to 15 fps at 33 percent scale. It says Veo 3.1 Lite is limited to 720p to 1080p, and it recommends MP4 with H.264 for compatibility.
Which steps does Sume cover?
The skill's point about matching frame rates also appears in Sume's Timeline docs: if you leave output.fps out, the render uses the rate the longest video source runs at, and a rate that differs from a source's means repeated or dropped frames, so motion can judder.
| Skill step | Sume surface | Note |
|---|---|---|
| Generate with Veo 3.1 | Video Router, a different catalog | Sume's router lists Gemini Omni Flash 1.1 and other models, not Veo ids |
| Concatenate clips | Timeline 1.0 | One audio spine plus 1 to 200 ordered video slots into one MP4 |
| Cut a range | Video trim | One clip, a start and one of end or duration |
| Crop, dim, filtergraph | Video filter | Allowlisted ops; an unbilled check route exists |
| Overlay a logo | Not in the docs I read | Video filter is described as filters only; confirm before relying on it |
| Video to GIF | Not in the docs I read | No GIF output is documented |
| Check media first | Video inspect | Probe and stills are unbilled |
How would an agent chain them on Sume?
Over the hosted MCP server an agent can call generate_video, then video_trim, video_filter and timeline_create, waiting on each with jobs_wait. Paid calls take an idempotency_key, and jobs_wait holds at most 55 seconds per call, so on wait_slice_expired you call it again with the same ids rather than resubmitting.
Timeline has a free plan route, POST /v1/timeline-1.0/plan, that returns the planned duration and an estimated cost before you render. The public rate is listed as $0.10 per output minute, rounded up, and the docs say to confirm the live rate in the catalog.
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio": {"mode": "silence", "duration_seconds": 16},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4", "start": 0, "duration": 8},
{"source_url": "https://media.sume.com/artifacts/artf_demo/b.mp4", "start": 0, "duration": 8}
]
}'Which should I choose?
If you want Veo 3.1 itself and run on Google Cloud, Google's skill is the direct path. Mind the date: Google's deprecations page lists the Veo 3.1 preview ids with a shutdown on October 22, 2026, and names gemini-omni-1.1-flash as the replacement. If you want one set of jobs with job ids, retries and a billing record, Sume's separate steps are the closer fit, with the gaps in the table.
What should I watch for when I merge clips?
Google's skill says clips need matching dimensions and frame rates before they are merged. That is good advice anywhere. Before you build a Timeline, probe each clip with Video inspect, which returns probe facts and stills without billing, and write down width, height and frame rate. Clips from different sources often run at different frame rates, for example 24 and 30, so a mixed edit will repeat or drop frames unless you set output.fps on purpose.
Timeline's default output is 1080 by 1920, which suits Shorts. A 16:9 source in a vertical timeline needs a crop or a pad decision first, and Video filter can crop one clip with an allowlisted op. Check the program with the unbilled POST /v1/video-filter/check before you pay for the encode, since a program that passes the check can still fail on the box with a structured job error.
Sources
Related posts
More in Comparisons
- gpt-4o-mini-tts instructions vs Sume TTS speed, volume, emotion
OpenAI steers gpt-4o-mini-tts with free-text instructions. Sume TTS exposes speed 0.6-1.5, volume 0.5-2 and a short emotion string. How the controls compare.
- GPT-6 Sol vs Luna vs Astra: context, prices, cutoffs
GPT-6 Astra, Sol and Luna share a 1,050,000-token window but differ 100x in price. A table from OpenAI's own pages, plus which one Sume runs Formats on.
- gpt-image-2.5 partial image streaming vs Sume's 400
OpenAI streams 0 to 3 partial images for gpt-image-2.5. Sume returns 400 streaming_not_supported for stream: true. What to send instead: a job and a poll.
- gpt-image-2.5 rate tiers (5 to 250 IPM) vs Sume plan queue
OpenAI limits gpt-image-2.5 Flare from 5 images per minute at Tier 1 to 250 at Tier 5. Sume limits processing concurrency by plan and queues the rest.
Written by Sume