Gemini app locks the chat while a video renders: run several at once

Gemini says a video takes a few minutes and you can't use that chat meanwhile. On Sume, valid jobs queue by plan: 1 at a time on Free, 4 on Pro, 20 on Scale.

5 min readSume
All posts

Google's Gemini app help page says a video takes a few minutes to generate and that you cannot interact with the same chat while it does, though you can start a new chat. If you want several variations running together, the app's answer is more chats. On Sume the answer is jobs: you submit them all, they queue, and your workspace plan sets how many process at once.

This page reads the Gemini Apps Help page on 2026-10-02 and Sume's Generation admission and Jobs and results docs.

What does the app do while a video renders?

Per the help page, the chat that is generating is blocked until it finishes, and a new chat is the way to keep working. The page says nothing about how many videos can render at the same time across chats, so do not plan a batch on a number. It also lists requirements that shape who can run anything at all: a Google AI plan or a qualifying Workspace license, being 18 or over, and being signed in.

What happens when I submit five jobs on Sume?

If the queue is full, new paid submissions fail with 429 queue_full. Request-rate limits give 429 rate_limited, and an empty balance gives 402 insufficient_credits before any provider work begins.

Sume generation concurrency by plan (read 2026-10-02)
PlanProcessing at onceQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

How do I submit a wave and wait for it?

Give every submit its own idempotency key so a retry cannot bill twice, store each job id, and poll with backoff. queued is a normal state, not a failure. Over the hosted MCP server, jobs_wait takes 1 to 20 job ids with wait_for set to all or any, and each call holds at most 55 seconds, so you repeat it on wait_slice_expired rather than asking for a longer wait.

The loop below submits three 9:16 variations of one prompt, then reads the status route for the first.

for i in 1 2 3; do
  curl -s -X POST https://api.sume.com/v1/video-router/generate \
    -H "Authorization: Bearer $SUME_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: wave-variation-$i" \
    -d '{"model":"gemini-omni-flash-1.1","prompt":"A paper boat crosses a puddle, variation '$i'. Soft rain sound.","duration":6,"resolution":"360p","aspect_ratio":"9:16","mode":"async"}'
  echo
done

curl https://api.sume.com/v1/jobs/job_123/status \
  -H "Authorization: Bearer $SUME_API_KEY"

What does this not fix?

Use the app when you are making one clip and can wait a few minutes. Use jobs when you want a set of options to choose from, and a record of what ran.

  • Queued is not faster. On the Free plan three jobs still run one after another.
  • Each job bills on its own. Use 360p drafts before you commit to 1080p.
  • One call makes one video. Variations are separate requests with separate keys.
  • Sume offers no chat for the clip; follow-up edits are new jobs with a video_url source.

Can I see what a failed job did?

Yes. A job ends as completed, failed or canceled, and a failed one carries a public error you can read from the status route. Events give a public timeline for debugging and recovery. Read those before you resubmit, and when you must resubmit, reuse the same idempotency key so the retry maps to the original request.

Compare that with the app, where a stuck video leaves you with a locked chat and a new chat as the only remedy. A job record is the advantage of an API here, not speed.

How should I size a wave for my plan?

Start from accepted job capacity, not from concurrency. On Free that is 6 jobs, so a seventh paid submission is rejected with queue_full until something finishes. On Pro it is 24, on Startup 48 and on Scale 120. A wave of 20 variations is fine on Pro, but it will take about five rounds of the processing slots, so plan the time as well as the spend.

Wait with the smallest call that answers your question. jobs_wait with wait_for: any returns as soon as one job is terminal, which is handy when you only need the first good clip. wait_for: all suits a wave you will review together. Pass include_results: true to get each finished job's result in the same answer, and read anything the answer omits with one batch jobs_result.

Never resubmit a paid create because a wait timed out. A 524 on a wait is a transport failure, not a job outcome, so re-issue the wait on the same ids.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume