Flow keeps 3 edit turns and History; build the audit trail on Sume
Flow's Omni edits hold context for up to 3 conversational turns and list every version in History. On Sume, each edit is a job id you can log and replay.

Google Flow lets you keep editing an Omni clip for up to 3 conversational turns without losing context, and its History panel lists all previous versions of the video. Sume has no conversation: every edit is a separate job with its own id, so the audit trail is whatever you record, and the cheap way to record it is the job id plus the idempotency key.
What exactly does Flow remember?
The Flow help page says you can continue for up to 3 conversational turns without losing the context, and that in the History panel you can find all previous versions of the video. It does not say what happens on turn four, so plan as if context ends there and re-describe the scene when you start a fresh chain.
The Gemini API describes the same idea as conversational editing through the Interactions API, where each instruction builds on the previous one. Sume's docs describe no such session, which is covered in one Sume job per edit.
How do I recreate the History panel on Sume?
Every Sume submit returns a job and a durable result. The Video Router edit takes a video_url source; the finished job's result carries the new video, which becomes the video_url of your next edit. Nothing else links version 2 to version 1, so write that link down yourself.
A simple ledger needs four fields per row: the Idempotency-Key you sent, the job id returned, the source URL, and the prompt. The Jobs and results page covers polling and the job status and result paths.
| Need | Flow | Sume |
|---|---|---|
| See earlier versions | History panel lists all previous versions | Your ledger; each result is its own durable file |
| See earlier prompts | Kept alongside versions, per the help page | The prompt you stored with the job id |
| Safe retry of one edit | Not described | Same Idempotency-Key replays the original job |
| Context across edits | Up to 3 conversational turns | None; restate the edit each time |
What does a three-edit chain look like?
Each call below uses its own key so a retry cannot create a duplicate. Paste the result URL from step one into step two.
Because an edit prompt has no memory of the earlier one, a later prompt that says only "make it darker" is ambiguous. Name the subject again: "Keep the same scene. Make the room darker." The Omni prompt-style notes in the Omni edit prompts post cover wording.
for n in 1 2 3; do
curl -s -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: chain-demo-edit-$n" \
-d "{\"model\":\"gemini-omni-flash-1.1\",\"prompt\":\"Edit $n: warm the lighting. Keep everything else the same.\",\"video_url\":\"$SRC_URL\",\"mode\":\"async\"}"
echo
doneWhat is the catch?
The loop above sends the same source three times, so it makes three parallel versions, not a chain. To chain, wait for each job, read its result URL, and use it as SRC_URL for the next. That is the real difference from Flow: its turns are a conversation, yours is a loop you own.
Re-encoding also compounds across a chain. Each generation is a new file, so check the clip with video inspect after the last step rather than assuming it matches the first.
Sources
Related posts
More in Media tools
- Flow Scenebuilder vs the Sume timeline API: sequence clips in code
Flow's Scenebuilder arranges clips in order. The Sume Timeline 1.0 route does the same from JSON: starts, fades, an audio spine, one MP4 out.
- Flow video edit limits: 60 s, 1 GB, 10 s window, and the Sume route
Flow edit and refine takes uploads up to 60 s and 1 GB but edits a 10 s window. Prepare the same clip for Sume's Gemini Omni video_url edit.
- FLUX 21:9 ultrawide image: FLUX.2 on Sume vs FLUX 3 ratios
FLUX 3 Image lists ratios from 21:9 to 9:21. On Sume, FLUX.2 Pro and Flex take 21:9 and 9:21 too, from a 13-ratio list with no auto. A banner request in Python.
- ImageKit video smart crop fo-face vs a Sume 9:16 reframe
ImageKit can track faces with fo-face when cropping video. Sume has no subject tracking: you choose a static crop, a fit mode, or one crop per shot.
Written by Sume