Veo upscaling on Vertex works on any video: Sume's 1.1x to 4x job
Google's Veo upscaling on Vertex AI is a private preview that takes any source video up to 4K. Sume's upscale job takes a public HTTPS video, 1.1x to 4x.

Google says its new Veo upscaling capability on Vertex AI is a standalone feature that raises video to 1080p or 4K and works on any source: Veo clips, other AI models or camera footage. It was in private preview when announced, with a public preview listed as coming soon. Sume has its own video upscale job, with a 1.1x to 4x scale ratio and a spend reservation that tops out at 30 seconds of input.
Google's facts come from its Cloud blog post Veo 3.1 Lite and a new Veo upscaling capability on Vertex AI, read on 2026-10-02. Sume's come from the API reference OpenAPI schema for Video Upscale 1.0.
What did Google announce?
The post is dated April 3, 2026 and covers two things. The first is Veo 3.1 Lite, the cost-focused tier of the Veo 3.1 family alongside Veo 3.1 and Veo 3.1 Fast, all with native audio on Vertex AI. The second is upscaling: standalone, separate from generation, up to 1080p or 4K, and usable on video from any source. Access is through the Vertex AI API and Vertex AI Media Studio. The post gives no price for upscaling and points to the Vertex pricing page.
Because the upscaler is separate from generation, it answers a real question: can I sharpen a clip a different model made? Google's answer is yes, if you are in the preview.
What does Sume's upscale job take?
The Video Upscale 1.0 request needs a public HTTPS video_url. Optional fields are scale_ratio from 1.1 to 4, with a default of 2, upscale_factor as an alias, duration_seconds from 1 to 30 for the spend reservation, and enhancement_tier of fast, standard or pro, defaulting to fast. If you omit duration_seconds, the reservation assumes 5 seconds, and the maximum is 30.
It is a normal Sume job: it returns a job id, you poll the status route, and a paid retry should reuse its idempotency key. Over the hosted MCP server the tool is video_upscale_create.
| Question | Google, Vertex AI | Sume Video Upscale 1.0 |
|---|---|---|
| Status | Private preview at launch; public preview coming soon | Documented route, no preview label |
| Source video | Any: Veo, other models, camera | Any public HTTPS video URL |
| Target | Up to 1080p or 4K | Scale ratio 1.1 to 4 |
| Input length | Not stated in the post | Reservation capped at 30 seconds |
| Quality control | Not stated | Tier: fast, standard or pro |
How do I call it?
The request below doubles a short clip with the middle tier. Keep the clip at 30 seconds or less; for a longer video, cut it first with Video trim and upscale the pieces.
curl -X POST https://api.sume.com/v1/video-upscale-1.0/upscale \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: upscale-2x-001" \
-d '{
"video_url": "https://example.com/clip.mp4",
"scale_ratio": 2,
"duration_seconds": 8,
"enhancement_tier": "standard"
}'Do I need an upscaler at all?
Maybe not. Sume's Gemini Omni Flash 1.1 already lists 360p, 720p, 1080p and 4K as resolution values, so for that model you can ask for the size directly. An upscaler earns its place when the source is a clip you already have, such as footage from another model or a camera, and a re-render is not an option. Neither service promises that an upscale adds detail that was never there, so check a frame before you ship it.
What can go wrong with any upscaler?
Upscalers raise pixel counts; they cannot restore what was never recorded. Faces, text and fine patterns can come out smoother or subtly changed. Check a few frames at full size before you ship, and keep the original.
Length is the practical limit on Sume: the catalog reserves at most 30 seconds of input per job, so plan on clips of 30 seconds or less. Cut a longer video with Video trim, upscale each piece, then join the pieces with Timeline. Match the frame rate across the pieces, or the join will repeat or drop frames.
Sume reserves spend from duration_seconds when you submit. If you leave it out, the reservation assumes 5 seconds, so set it to the real length of the clip. Check the live price in the catalog, because rates change.
Sources
Related posts
More in Media tools
- Check a video filter program for free before you encode it
POST /v1/video-filter/check runs the same schema and allowlist as the encode and returns diagnostics without creating a job. Encoding is $0.02 a job.
- Video frames vs video inspect: max_edge 16 vs 64 and the 768 default
Sume video-frames max_edge is 16 to 2160 and keeps source size if omitted. Video-inspect max_edge is 64 to 2160 and defaults to 768.
- inspect_source_has_no_audio: check probe.has_audio before transcribe
transcribe true on a silent clip fails as inspect_source_has_no_audio. Probe first with frames false and read probe.has_audio, then ask for the transcript.
- Which AI video models take reference images in Sume's Videos panel?
Auto, Kling 3.0, Wan 3.0 and MiniMax H3 show reference slots in Sume's panel; H3 Max and Grok do not. Plus the API limits for image, video and audio references.
Written by Sume