Vidu Q4 has no text-to-video: Sume rows that take a plain prompt
Vidu's Q4 page lists only image-to-video and reference-to-video. Sume has no Vidu row, but several rows start from a text prompt alone. Compare limits.

No: the Vidu Q4 page lists two modes only, Image-to-Video and Reference-to-Video, so a text-only prompt has no Q4 mode to run in (read 2026-10-10). Sume does not carry a Vidu row either, but its catalog has several rows that accept a text prompt with no image: Seedance 2.5, Wan 3.0, Kling 3, MiniMax H3, and Gemini Omni Flash 1.1.
This matters if you came to Vidu Q4 for the new 4K and 16-second limits and your workflow starts from a script, not a picture. You then need either a picture step in front of Q4, or a different model. This post covers the second path.
What Vidu Q4 accepts, as published
The Vidu page (read 2026-10-10) gives these limits. Reference-to-video takes 1 to 16 seconds. Image-to-video takes 3 to 16 seconds. Output goes up to 4K. You can pass up to 3 voice references and between 1 and 15 images. The page shows no price, so this post does not price Q4.
A prompt-only call is the one thing the page does not list. If your pipeline sends only text, check Vidu's own API reference before you plan around Q4.
| Mode | Duration | Inputs |
|---|---|---|
| Image-to-Video | 3 to 16 s | One start image plus prompt |
| Reference-to-Video | 1 to 16 s | 1 to 15 images, up to 3 voice references |
| Text-to-Video | Not listed | Not listed |
Sume rows that start from text alone
Sume's Video Generation docs describe POST /v1/videos as text-to-video with optional reference images. The registry in the Sume API code marks these rows as text-capable. The limits below come from the docs and the Video Router capability table in the code.
Veo and Tencent rows are also text-to-video, but they only appear in a workspace when GET /v1/videos/models returns their ids, so check the catalog before you plan around them.
| Sume id | Duration | Resolutions | Audio |
|---|---|---|---|
| seedance-2.5 | 4 to 30 s | 480p, 720p, 1080p | Yes |
| wan-3.0 | 2 to 30 s | 480p, 720p, 1080p | Yes |
| kling-3 | 4 to 15 s | 1080p only | Yes |
| minimax-h3 | 5 to 15 s | 480p, 768p (2K and 4K billed if asked) | Yes |
| gemini-omni-flash-1.1 | 3 to 10 s | 360p, 720p, 1080p, 4K | Yes |
Longest shot and cheapest start
If you need the longest single shot, Seedance 2.5 and Wan 3.0 both go to 30 seconds, which is almost double Q4's 16. If you need the cheapest draft, Wan 3.0 at 480p bills $0.0625 per second on Sume, so a 5-second draft is $0.3125. A 5-second Seedance 2.5 draft at 480p bills $1.343385.
Those numbers are the provider list price times 1.25, applied at submit, as the Video Generation docs describe in the billing row. The estimator output already includes the 1.25 factor, so do not multiply again.
When to still use Q4
Q4 is the better fit when you already hold a strong start image and want up to 16 seconds with native audio and voice references. Sume cannot offer that mix today, because it has no Vidu row.
For text-first work, pick a row above, send model and prompt, and poll the job. Read the catalog first with GET /v1/videos/models, since each row lists its own durations and resolutions, and Sume rejects size, seed and provider.options with a 400.
- Text script only: use Seedance 2.5 or Wan 3.0 on Sume.
- Start image in hand and 16 s wanted: Vidu Q4, outside Sume.
- Need a voice reference: see the linked audio-reference post.
Checking the catalog before you build
Do not copy the table above into code. Call GET /v1/videos/models with your key and read supported_durations, supported_resolutions and supported_input_references for each id. Those three fields tell you in one response whether a row takes your prompt, your length and your reference files.
Sume describes this endpoint as the place to find the supported models, their capabilities and their prices, and the Video Generation page shows the exact response shape. A row that is missing from the list is simply not available to your workspace.
Sources
Related posts
More in Comparisons
- Vidu Q4 takes 15 reference images; Sume rows cap at 9 or 10
Vidu Q4 accepts 1 to 15 reference images. Sume has no Vidu row: most rows allow 9 images, Wan 3.0 and Omni allow 10. Caps for images, video and audio refs.
- Vidu Q4 voice references vs Sume's reference audio inputs
Vidu Q4 takes up to 3 voice references. Sume has no Vidu row; Seedance, Wan 3.0 and MiniMax H3 accept audio references, but voice cloning is app-only, not API.
- Voxtral Mini 4B Realtime Arabic: open weights vs a Sume STT job
Mistral's Arabic streaming model arrived Oct 8 as Apache 2.0 weights, not a service. For Arabic transcripts without hosting, Sume STT takes language_code ar.
- Wan 3.0 edit and extension at Alibaba vs Sume's Wan row
Alibaba lists Wan 3.0 editing and extension; Sume's wan-3.0 row does text, image, end frame and references only. Edit with Omni Flash 1.1 or H3 Max Recast.
Written by Sume