Gemini Omni previous_interaction_id vs a Sume job per request
Google chains Omni 1.1 Flash extensions with previous_interaction_id. Sume has no such field: each request is its own job, and an edit takes a video_url.

In Google's Gemini API, previous_interaction_id links a new client.interactions.create call to an earlier video interaction so the model continues that scene. Sume's published video contract has no equivalent field: every request is a separate job with its own id, and you carry state yourself by passing media back in.
The Google side is from its Omni 1.1 Flash developer post; the Sume side is from Video generation and Video Router, both read 2026-10-01.
What does previous_interaction_id do in the Gemini API?
Google's example calls client.interactions.create with previous_interaction_id set to an earlier video interaction and a text input of "Continue the scene." Google says Omni 1.1 can analyze up to 10 seconds of prior context, where earlier models referenced only the final second, and that you extend in 10-second increments up to 40 seconds in total.
The state, in other words, lives on Google's side and is addressed by an interaction id.
How does Sume handle the same job?
Sume does not chain by id. A request to the Video Router with gemini-omni-flash-1.1 returns a job; poll it and read the result when it completes. To build on a result, you send the next request yourself, using the clip as input. The docs list Omni Flash 1.1 at 3–10 seconds at 360p/720p/1080p/4K in 16:9 or 9:16 with native synced audio, and it exposes video edit mode through the Video Router video_url field.
That edit mode is the closest Sume analogue to "take this clip and change it". For longer sequences, see the 40-second extension comparison.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-edit-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Replace the bottle with an apple. Keep everything else the same.",
"video_url": "https://example.com/clip.mp4",
"resolution": "720p",
"mode": "async"
}'What changes in my code?
| Concern | Gemini API (Google post) | Sume |
|---|---|---|
| Continuation | previous_interaction_id | Next request carries the clip as video_url (edit) or other inputs |
| Unit of work | An interaction | A job with its own id |
| Result | Returned by the interaction | The job result, read after the job completes |
| Safe retry | Not covered in the post | Idempotency-Key makes retries safe; a replay returns the original job |
Why does the Idempotency-Key matter here?
Because each step is its own paid job, a retried network call could otherwise create a second one. The docs say to send Idempotency-Key so retries are safe. Store one key per logical step, and see idempotency keys for AI video APIs.
Sources
Related posts
More in Developers
- Text to speech mulaw 8000 Hz: Gemini 3.8 and Sume TTS
Sume TTS can return pcm_mulaw or pcm_alaw at 8000 Hz in wav or raw containers. Gemini 3.8 TTS does it with audio/mulaw and audio/alaw mime types.
- Gemini TTS voice design voice_ id vs Sume voice ids
Gemini voice design returns a persistent voice_ id from a text prompt. Sume TTS accepts only a voice UUID or voi_ library id, and rejects other shapes with 400.
- GPT Image 2 input_fidelity: omit it; Sume returns 400
OpenAI says to omit input_fidelity for gpt-image-2 because inputs run at high fidelity. Sume lists no such field and rejects unlisted parameters with 400.
- GPT Image moderation_blocked vs Sume content_policy_rejected
OpenAI returns moderation_blocked with moderation_details. On Sume, policy refusals are grouped under content_policy_rejected. What to read and when to retry.
Written by Sume