Gemini CLI MCP timeout is 600000 ms; Sume jobs_wait caps at 55 s
Gemini CLI waits up to 10 minutes per MCP request, but Sume cuts jobs_wait at 55 seconds. Set the client timeout low and loop on wait_slice_expired.

Gemini CLI's per-server timeout defaults to 600,000 ms, which is 10 minutes. Sume's remote jobs_wait never holds a request longer than 55 seconds, so the client limit is never the one that fires. Set it to about 60,000 ms so a dead connection is reported in a minute, and have the agent repeat jobs_wait for long renders.
| Clock | Value |
|---|---|
Gemini CLI timeout default | 600,000 ms |
Sume jobs_wait default | 50 s |
Sume jobs_wait cap | 55 s |
Sume script_run timeout_seconds | 5 to 55 |
What goes wrong with a long client timeout
The Sume docs say that a wait is a single HTTP request that stays open with no data moving, and that edges close such requests. A 524, 522, 523 or 525 on jobs_wait is a transport failure, not a job result. With a 10-minute client timeout, the agent may sit on a dead call far longer than the slice it asked for.
Config
Lower the timeout on the Sume entry only. Other servers keep their own values.
{
"mcpServers": {
"sume": {
"httpUrl": "https://mcp.sume.com/mcp",
"timeout": 60000
}
}
}The loop to give the agent
Ten-minute renders are normal for some video work. Do not ask for a ten-minute wait; wait again.
- Submit once and keep the job id.
- Call
jobs_waitwith that id. - If the answer says
wait_slice_expired, call it again with the same id. - Never re-submit the paid create after a timeout. The job keeps running and billing.
- For several jobs, pass up to 20 in
job_idsinstead of one call per job.
Why 60 seconds
Sixty thousand milliseconds is five seconds above the Sume cap of 55. It lets a normal slice complete and still reports a hung call quickly. If you use script_run, remember that timeout_seconds goes up to 55 as well, so the same limit covers it.
What a long timeout hides
With the default of ten minutes, a call that the edge dropped does not fail for a long time. The agent sits idle, the job continues to bill, and you may start a second job meanwhile. A short client limit turns the silent hang into a prompt retry of the wait, which is cheap and safe.
If you must run very long scripts, move them off the wait path. Submit the work, store the id, and let a separate scheduled check call jobs_status or a single jobs_wait later. That pattern keeps every individual tool call short and makes each one easy to retry. It also fits the way Sume bills: the job runs and bills whether or not any client is waiting on it, so there is no gain in keeping a connection open.
Before you rely on this setup, run a short acceptance test with a read-only credential. Connect, call mcp_health, call tools_list, and read one job with jobs_status. Record the tool count you see, so you can notice later if a credential change alters it. Then repeat the test after any config edit. A five-minute test like this catches most wiring mistakes before they cost money, and it gives you a baseline to compare against when something behaves differently next week.
Sources
Related posts
More in Integrations
- Can GLM-5.3 call Sume's hosted MCP? Function calling, not MCP
Z.ai's GLM-5.3 page lists function calling and does not mention MCP. How a GLM agent reaches Sume through an MCP client or through plain HTTP.
- Kling lists MCP and CLI tools for batch video; Sume's hosted MCP
Kling's home page mentions MCP and CLI tools for batch generation. Sume's hosted MCP at mcp.sume.com has generate_video; its Kling row is kling-3 only.
- LinkedIn wants a text-only SRT; Sume burns captions, so build one
LinkedIn video ad captions must be a text-only SRT. Sume burns captions into pixels and has no SRT export. Build the file from video-inspect segments.
- Make takes 300 webhook requests per 10 seconds: a Sume queue fits
Make accepts 300 incoming requests per 10 seconds. A 100-item Sume queue with per-item webhooks and 16 in flight stays far below that. When 429 and 400 matter.
Written by Sume