Google Vids point-in-time editing: timed text via Sume cues
Google Vids now shows only what is on screen at the playhead. Sume has no canvas, but caption cues with start and end place timed text by API.

Google Vids now shows only the elements that are on screen at the playhead's timestamp, so overlapping text boxes, stickers and images stop piling up on the canvas. Sume has no editing canvas to compare it with. The nearest Sume habit is to declare each piece of timed text up front as a cue with a start and an end, and let a caption job burn it in at exactly those seconds.
This post compares the two ideas using Google's own announcement, Point-in-time editing now available in Google Vids (read 2026-10-10), and Sume's Video captions reference. It does not claim that Sume matches the Vids editor: it is an API, and it covers only text.
What Google shipped in Vids
Google describes a synchronized canvas and timeline. In its words, when a user is editing, they see exactly what appears in the video at that timestamp, without the clutter of overlapping text boxes, stickers and images. The page says it is enabled by default, needs no administrator action and has no user setting.
The rollout dates on that page are the useful part for planning. Rapid Release domains began a gradual rollout on September 29, 2026, taking up to 15 days. Scheduled Release domains start a full rollout on October 13, 2026, taking 1 to 3 days.
| Item | What the page says |
|---|---|
| What it does | The canvas shows only what is on screen at the selected timestamp |
| Rapid Release domains | Gradual rollout began September 29, 2026, up to 15 days |
| Scheduled Release domains | Full rollout starts October 13, 2026, 1 to 3 days |
| Admin action | None needed |
| User setting | None; on by default |
The Sume equivalent: cues with a start and an end
A Vids editor drags a text box along a timeline. A Sume caller lists the same text as data. The video-captions job accepts cues (or segments), each with text, start and end in seconds, and burns that text at those times without running speech-to-text. The docs call this the path for silent clips, but nothing stops you using it on a clip with speech when you want authored overlay text instead of a transcript.
Because a cue exists only between its start and end, you get the same result Google's canvas now previews: at any second, only the cards whose window contains that second are on screen. You send only one of script_text, words, cues and segments per request, so a job is either transcript-driven or authored.
video_urlmust be a public HTTPS video URL that Sume can fetch; localhost, private-network, non-HTTPS and signed URLs are rejected.stylepicks the look;designoverrides colors, typography, placement, phrasing and motion for one request, except onpunchandtiktok-green, which do not supportdesign.design.placement.anchor_ratiosets the center of the line as a fraction of the frame height, which is how you keep two cards from landing on the same spot.- The price is $0.20 per accepted standalone caption job for videos of up to 60 seconds, under the current fixed estimate; the live price is in
GET /v1/catalog.
A request that places two timed cards
The sample sends two cards on one clip. It sets an Idempotency-Key, so a retry does not queue a second paid job. The response is a job; poll it with GET /v1/jobs/:id/status and read the captioned video_url from the result, as in the Video captions page.
import asyncio, os
import httpx
CUES = [
{"text": "Free shipping", "start": 0.5, "end": 2.5},
{"text": "Ends Sunday", "start": 3.0, "end": 5.0},
]
async def main():
key = os.environ["SUME_API_KEY"]
async with httpx.AsyncClient(base_url="https://api.sume.com", timeout=60) as c:
r = await c.post(
"/v1/video-captions",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": "timed-text-001"},
json={"video_url": "https://example.com/clean.mp4", "style": "black-outline", "cues": CUES},
)
print(r.status_code, r.json())
asyncio.run(main())Check the result at the cue times
Google's canvas lets an editor scrub and look. With an API you check by sampling. If the captioned video_url is on media.sume.com, Video frames can pull stills at times you name: at takes 1 to 24 values, each in the range 0 up to the clip duration. Ask for one second inside each cue and one second between them, and you can see that the card is present in the first and gone in the second.
Video frames is billed by its Modal compute rather than at a flat rate, and a submit always returns 202, so poll it. The frames come back as durable image artifacts at source size unless you set max_edge.
What Sume does not do here
There is no stickers-at-a-timestamp feature in the caption job, and no canvas. Cues carry text only. A still image can go over a video through Timeline compose, but docs say the still stays on screen for the whole clip, so it is not a timed element.
If the question is how to arrange many timed elements by hand, Vids is the right tool and the new view helps. If the question is how to stamp the same timed text onto hundreds of clips, the cue list is data you can generate, and a job per clip costs the same flat fee.
| Need | Google Vids | Sume video-captions |
|---|---|---|
| See only what is on screen now | Yes, by default | Not applicable; sample stills instead |
| Timed text | Text boxes on the timeline | cues with start and end |
| Timed stickers or images | Shown on the canvas | Not supported by captions |
| Repeat across many clips | Manual per project | One request per clip |
Sources
Related posts
More in Comparisons
- Grok Imagine's 4 keyframes and 7 references vs Sume's one image
xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.
- Grok Imagine Lite upscales to 1080p: what Sume offers instead
xAI describes Grok Imagine Video 1.5 Lite as lowest cost with upscaled 1080p. Sume has no Lite row; it lists grok-imagine-video-1.5 at a flat rate. Compare.
- Grok Imagine's 3 voice references vs Sume reference-audio rows
xAI's Grok Imagine 1.5 takes up to 3 voice references. Sume's Grok row takes none; Seedance, Wan 3.0 and MiniMax accept reference audio under limits.
- Grok Imagine Lite draft plus Sume upscale: what a 10 s clip totals
xAI lists Lite at $0.020 per second and 1.5 at $0.080. Add Sume video upscale at $0.009 per input second and a 10 s clip totals $0.29. Read 2026-10-10.
Written by Sume