Gemini 3.7 Flash now routes to 3.8: where a media client pins ids
Google auto-routes gemini-3.7-flash to gemini-3.8-flash since Oct 8. Where Sume lets you pin a model, where it routes for you, and what the job records.

On October 8, 2026 Google's release notes say Gemini 3.7 Flash is deprecated and requests are automatically routed to gemini-3.8-flash. Your code did not change, but the model behind the string did. If you generate media through an API, decide per call whether you want that behavior or a pinned id.
On Sume there are two explicit choices, plus one trap. You can pin a catalog id, you can ask for sume/auto, and a retired id may be accepted as an alias that runs a newer model.
Pin or route, per endpoint
For video, GET /v1/video-router/models lists the catalog and POST /v1/video-router/generate takes a model from it, for example seedance-2.5. The docs recommend POST /v1/videos with model: "sume/auto" when you want Sume to choose. Limits differ by model: Seedance 2.5 accepts 4-30 s, Wan 3.0 accepts 2-30 s, and the Auto controls default to 720p and 8 s with 3-10 s clips.
For images, the docs state that Nano Banana 2 is retired: google/nano-banana-2 and nano-banana-2 still work and run as Nano Banana 2.1, and the job stores the 2.1 id. So an old string keeps working, but your logs show a different model than the one you sent.
| You send | What runs | Where to confirm |
|---|---|---|
| Catalog id, e.g. seedance-2.5 | That model, within its limits | GET /v1/video-router/models |
| model: sume/auto on /v1/videos | Sume chooses; 3-10 s clips, default 720p and 8 s | POST /v1/videos docs |
| nano-banana-2 (retired) | Runs as nano-banana-2.1; job stores 2.1 id | The job record |
| Unknown id | 404 model_not_found | GET /v1/catalog |
What to do in a client
Treat the id you send and the id the job records as two fields. Log both. When they differ, you have an alias at work, and that is the signal to update your config.
Do not rely on an alias for cost control either. Prices are per model and resolution, so an alias that lands on a newer model can change the bill. Run dry_run or read the catalog before a large wave.
- Pin a catalog id in production jobs where output consistency matters.
- Use
sume/autoonly where any acceptable model will do. - Alert on
model_not_found, which is a 404 and not a retry case. - Re-read the catalog on a schedule instead of caching ids forever.
A config pattern
Keep model ids in one config file with two fields per use: the id you request and the id you expect the job to record. A nightly check submits a tiny job for each entry, reads the model from the job, and fails when the two differ. That catches silent aliasing, whether it is a vendor routing an old string to a newer model or Sume running a retired image id as its successor.
Pair it with a catalog check. GET /v1/catalog is public and needs no key, so the check can confirm that each pinned id still exists before you spend anything. A missing id then shows up as a failed check instead of a 404 model_not_found in production.
Sources
Related posts
More in Developers
- Gemini Deep Research agent shuts down Oct 23: an async alternative
Google marked deep-research-pro-preview-12-2025 for shutdown on Oct 23, 2026. What an async Sume Agent Completion run does, and what it does not do like it.
- Gemini Omni Flash 1.1 prompt length on Sume: a 20,000-character check
The Sume catalog constraint for Gemini Omni Flash 1.1 is a prompt of at most 20,000 characters, 3 to 10 seconds. A short Python preflight check before you pay.
- Omni Flash defaults to 720p and 16:9: set both on every request
Google's Omni docs default to 720p and 16:9; Sume's Auto defaults to 720p and 8 seconds. Why to send resolution, ratio and duration every time, with prices.
- Gemini Omni edit mode on Sume: video_url only, 720p, no aspect_ratio
Omni edit mode takes a prompt and one video_url. Resolution defaults to 720p, aspect_ratio is rejected, and no image or reference field may be added.
Written by Sume