Google Vids point-in-time editing: timed text via Sume cues

Google Vids now shows only what is on screen at the playhead. Sume has no canvas, but caption cues with start and end place timed text by API.

5 min readSume
All posts

Google Vids now shows only the elements that are on screen at the playhead's timestamp, so overlapping text boxes, stickers and images stop piling up on the canvas. Sume has no editing canvas to compare it with. The nearest Sume habit is to declare each piece of timed text up front as a cue with a start and an end, and let a caption job burn it in at exactly those seconds.

This post compares the two ideas using Google's own announcement, Point-in-time editing now available in Google Vids (read 2026-10-10), and Sume's Video captions reference. It does not claim that Sume matches the Vids editor: it is an API, and it covers only text.

What Google shipped in Vids

Google describes a synchronized canvas and timeline. In its words, when a user is editing, they see exactly what appears in the video at that timestamp, without the clutter of overlapping text boxes, stickers and images. The page says it is enabled by default, needs no administrator action and has no user setting.

The rollout dates on that page are the useful part for planning. Rapid Release domains began a gradual rollout on September 29, 2026, taking up to 15 days. Scheduled Release domains start a full rollout on October 13, 2026, taking 1 to 3 days.

Google Vids point-in-time editing, from Google's page (read 2026-10-10)
ItemWhat the page says
What it doesThe canvas shows only what is on screen at the selected timestamp
Rapid Release domainsGradual rollout began September 29, 2026, up to 15 days
Scheduled Release domainsFull rollout starts October 13, 2026, 1 to 3 days
Admin actionNone needed
User settingNone; on by default

The Sume equivalent: cues with a start and an end

A Vids editor drags a text box along a timeline. A Sume caller lists the same text as data. The video-captions job accepts cues (or segments), each with text, start and end in seconds, and burns that text at those times without running speech-to-text. The docs call this the path for silent clips, but nothing stops you using it on a clip with speech when you want authored overlay text instead of a transcript.

Because a cue exists only between its start and end, you get the same result Google's canvas now previews: at any second, only the cards whose window contains that second are on screen. You send only one of script_text, words, cues and segments per request, so a job is either transcript-driven or authored.

  • video_url must be a public HTTPS video URL that Sume can fetch; localhost, private-network, non-HTTPS and signed URLs are rejected.
  • style picks the look; design overrides colors, typography, placement, phrasing and motion for one request, except on punch and tiktok-green, which do not support design.
  • design.placement.anchor_ratio sets the center of the line as a fraction of the frame height, which is how you keep two cards from landing on the same spot.
  • The price is $0.20 per accepted standalone caption job for videos of up to 60 seconds, under the current fixed estimate; the live price is in GET /v1/catalog.

A request that places two timed cards

The sample sends two cards on one clip. It sets an Idempotency-Key, so a retry does not queue a second paid job. The response is a job; poll it with GET /v1/jobs/:id/status and read the captioned video_url from the result, as in the Video captions page.

import asyncio, os
import httpx

CUES = [
    {"text": "Free shipping", "start": 0.5, "end": 2.5},
    {"text": "Ends Sunday", "start": 3.0, "end": 5.0},
]

async def main():
    key = os.environ["SUME_API_KEY"]
    async with httpx.AsyncClient(base_url="https://api.sume.com", timeout=60) as c:
        r = await c.post(
            "/v1/video-captions",
            headers={"Authorization": f"Bearer {key}", "Idempotency-Key": "timed-text-001"},
            json={"video_url": "https://example.com/clean.mp4", "style": "black-outline", "cues": CUES},
        )
        print(r.status_code, r.json())

asyncio.run(main())

Check the result at the cue times

Google's canvas lets an editor scrub and look. With an API you check by sampling. If the captioned video_url is on media.sume.com, Video frames can pull stills at times you name: at takes 1 to 24 values, each in the range 0 up to the clip duration. Ask for one second inside each cue and one second between them, and you can see that the card is present in the first and gone in the second.

Video frames is billed by its Modal compute rather than at a flat rate, and a submit always returns 202, so poll it. The frames come back as durable image artifacts at source size unless you set max_edge.

What Sume does not do here

There is no stickers-at-a-timestamp feature in the caption job, and no canvas. Cues carry text only. A still image can go over a video through Timeline compose, but docs say the still stays on screen for the whole clip, so it is not a timed element.

If the question is how to arrange many timed elements by hand, Vids is the right tool and the new view helps. If the question is how to stamp the same timed text onto hundreds of clips, the cue list is data you can generate, and a job per clip costs the same flat fee.

Vids canvas versus a Sume cue list (Vids column read 2026-10-10; Sume column from the Video captions docs)
NeedGoogle VidsSume video-captions
See only what is on screen nowYes, by defaultNot applicable; sample stills instead
Timed textText boxes on the timelinecues with start and end
Timed stickers or imagesShown on the canvasNot supported by captions
Repeat across many clipsManual per projectOne request per clip

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume