Gemini Deep Research agent shuts down Oct 23: an async alternative
Google marked deep-research-pro-preview-12-2025 for shutdown on Oct 23, 2026. What an async Sume Agent Completion run does, and what it does not do like it.

Google's Gemini API release notes list deep-research-pro-preview-12-2025 for shutdown on October 23, 2026, 14 days after this post, and say two newer versions are available. The first thing to do is migrate to one of those. This post is for a different case: your job was never really 'research' and you used the agent because it was the only long-running task endpoint you had.
If that is you, Sume's Agent Completions endpoint is one candidate. It is not a drop-in for Deep Research and the docs make no such claim. It runs the Sume Studio Agent on an ad-hoc prompt, with a sandbox, tools, an MCP bridge and media generation.
What an Agent Completion is
POST /v1/agent/completions returns 202 and a receipt, never choices[]. You poll status_url until next_action is no longer poll_status, or you register communication.webhook_url and get one signed POST when the run completes or fails. There is no streaming and no synchronous wire today.
The request uses the OpenAI messages[] shape, with system and user turns only. You can send exactly one of instruction or messages. Each call runs in a new thread, so there is no continuation by thread_id.
| Item | Value |
|---|---|
| Create | POST /v1/agent/completions, returns 202 + agent.run receipt |
| Read | GET /v1/agent-runs/{id}, /status, /result |
| Cancel | POST /v1/agent-runs/{id}/cancel |
| Model field | Only sume-agent |
| Required cap | generation_spend_cap_usd, no default |
| Scopes | agent_completions:read, agent_completions:write |
| Streaming / sync | Not available |
Gaps to check before you move
Match your requirements against what is listed as not available: no streaming, no choices[], no assistant turns, no thread continuation, only image attachments (up to 30), and completions are user-owned, not team-owned. Service-account keys cannot create them and get 403 insufficient_scope.
Also check the spend rule. A backend caller has no approval prompt, so the cap replaces it. A request without generation_spend_cap_usd fails with 400 invalid_request. Set it to the most you accept for that one run, using the metered rates on the API pricing page.
- Keys created before Agent Completions shipped lack the scopes. Create a new key and rotate to it.
- Use
Idempotency-Keyon create. The same key with the same payload returns the original receipt withidempotency_hit: true. - For structured results, pass
output_schemaand readoutputfrom the completed receipt.
A migration order
Do the vendor migration first: update the Gemini model string before Oct 23. Then, if you decide the task is a general agent run, prototype it on Sume with a small cap and compare outputs on your own inputs. Only the comparison on your data can tell you whether the result is acceptable.
A small trial plan
Pick three real prompts from your current Deep Research usage and run each as an Agent Completion with a cap of a few dollars. Record the receipt usage, the time to completion, and whether output satisfies a schema you wrote. Use Idempotency-Key values like trial-1-prompt-a so a rerun does not bill twice.
Note that the cap is for generation spend on the run. Read the metered rates on the API pricing page before you pick it, and start low. If a run ends failed, the receipt keeps any artifacts and an output_error, which is useful when you compare outputs. Keep the vendor migration in place until the trial shows an acceptable result, because the Oct 23 date does not move.
Sources
Related posts
More in Developers
- Gemini Omni Flash 1.1 prompt length on Sume: a 20,000-character check
The Sume catalog constraint for Gemini Omni Flash 1.1 is a prompt of at most 20,000 characters, 3 to 10 seconds. A short Python preflight check before you pay.
- Omni Flash defaults to 720p and 16:9: set both on every request
Google's Omni docs default to 720p and 16:9; Sume's Auto defaults to 720p and 8 seconds. Why to send resolution, ratio and duration every time, with prices.
- Gemini Omni edit mode on Sume: video_url only, 720p, no aspect_ratio
Omni edit mode takes a prompt and one video_url. Resolution defaults to 720p, aspect_ratio is rejected, and no image or reference field may be added.
- Omni reference-to-video with 3 clips of 3 s: a 10 s output is $1.25
Gemini Omni Flash 1.1 takes up to 10 reference images and 3 reference clips of 3 seconds each. A 10-second 720p output is 10 x $0.125 = $1.25 on Sume.
Written by Sume