CapCut lists Seedance 2.5 and Gemini Omni. Does Sume's API carry them?
CapCut's tools page names several video and image models. Sume's API documents Seedance 2.5 and Gemini Omni Flash 1.1; the rest are not in its docs.

Two of the models CapCut names on its tools page are in Sume's documented API: Seedance 2.5 as seedance-2.5 and Gemini Omni as gemini-omni-flash-1.1. The others on that page, such as Kling, PixVerse, Veo, Happy Horse and Nano Banana Pro, do not appear in the Sume docs I read, so this post does not say Sume carries them.
CapCut's list is from its tools page, read 2026-10-10. The page lists Dreamina Seedance 2.5, Nano Banana Pro, Gemini Omni, Kling, PixVerse, VEO and Happy Horse, along with automatic captions, text to speech and filler-word removal. It describes an editing app; the page I read did not mention an API, so nothing below assumes one. Sume's side comes from the Videos API, the Video Router and the Images docs.
What the Sume docs name
The Videos API documents a request that takes a bare catalog model id, with sume/auto as a way to let Sume choose. The examples use seedance-2.5 and seedance-2, and the limits paragraph names wan-3.0, minimax-h3, minimax-h3-max and gemini-omni-flash-1.1. The authoritative list is the live endpoint GET /v1/video-router/models, which also reports each model's capabilities, so read it before you code against a model.
For images, the docs list bare ids such as gpt-image-2 and nano-banana-2.1 as aliases of their org/slug forms, and state that Nano Banana 2 is retired and now runs as 2.1. Nano Banana Pro does not appear on the page I read.
| Name on CapCut's page | Documented in Sume's API docs? | What the docs say |
|---|---|---|
| Dreamina Seedance 2.5 | Yes, as seedance-2.5 | 4 to 30 seconds at 480p, 720p and 1080p |
| Gemini Omni | Yes, as gemini-omni-flash-1.1 | 3 to 10 seconds, 360p to 4K, 16:9 or 9:16, native synced audio |
| Nano Banana Pro | Not in the pages read | The docs name nano-banana-2.1, not a Pro variant |
| Kling | Not in the pages read | No claim made |
| PixVerse | Not in the pages read | No claim made |
| VEO | Not in the pages read | No claim made |
| Happy Horse | Not in the pages read | No claim made |
Limits that differ from what you might assume
Check duration before you port a prompt. Seedance 2.5 and Wan 3.0 accept up to 30 seconds in the Sume docs, and most other catalog models stop at 15. Gemini Omni Flash 1.1 is the exception in the other direction: it accepts 3 to 10 seconds, so a 12 second request is outside its range. Its edit mode uses the Video Router video_url field, and in that mode you do not send aspect_ratio or duration.
Pricing also differs from a consumer app. The Video Router doc says Sume bills the list price times 1.25 for each model, and Gemini Omni is billed per output second as a function of resolution. This post does not quote CapCut's own pricing, because the page I read did not give it.
When CapCut is the better tool
If you want to open a clip, trim by eye, add captions and publish from a phone, CapCut's list describes an editor built for that, and an API is the wrong shape. Sume fits when a program makes the clip: a request goes in, an asynchronous job comes out, and you poll or take a webhook. The two can sit together: generate through the API, then edit by hand in the app.
A practical way to decide is to write down the model you actually need. If it is Seedance 2.5 or Gemini Omni, send a request to POST /v1/videos with that id and read the capabilities first. If it is anything else on CapCut's list, ask for it in the model's own vendor documents, because Sume's docs do not promise it.
Sources
Related posts
More in Comparisons
- ChatGPT Image 2 vs 2.5 on Sume: $0.26375 vs $0.065875 and what differs
On Sume, GPT Image 2.5 high quality at 1024 costs $0.065875, a quarter of GPT Image 2 at $0.26375, and adds mask_url, background and 16 references.
- Creatomate RenderScript vs a Sume Timeline document: field map
Creatomate's RenderScript is a general scene JSON; Sume's Timeline 1.0 is one audio spine plus video slots. Field-by-field map and what Sume cannot express.
- Creatomate template modifications vs a Sume Format run input
Creatomate fills a fixed template through modifications; a Sume Format run takes free-form input and generates new media. Request shapes and when each fits.
- D-ID audio talks: 15 MB, 5-10 minutes vs Sume's 4-60 s script
D-ID takes up to 40,000 characters of text or 15 MB of audio for a talk. Sume takes an English script of 4 to 60 seconds. Which input fits (read 2026-10-10).
Written by Sume