Replicate allow_fallback_model: Nano Banana Pro vs Sume
Replicate can fall back from Nano Banana Pro to Seedream 5.0 lite and bills the fallback. Sume's allow_fallbacks is accepted but has no effect. What to do.

Replicate's Nano Banana Pro can fall back to ByteDance Seedream 5.0 lite when Google's API is at capacity, but only if you set allow_fallback_model to true, and you are charged the fallback model's cost. Sume has no equivalent: its image API accepts provider.allow_fallbacks, but with one sume endpoint per model the flag has no effect, so a capacity problem shows up as a queued job or an error, not a different model.
Replicate details come from its changelog entry of March 2, 2026, read on 2026-10-03. Sume details come from the image generation docs, errors and credits, and generation admission.
What does Replicate's fallback actually do?
Per the changelog, setting allow_fallback_model to true lets Nano Banana Pro route to Seedream 5.0 lite when Google's API reaches capacity. The entry states: "You're charged the cost of the fallback model, not Nano Banana Pro."
It also lists limits where the fallback cannot help: Seedream 5.0 lite does not support 1K (it generates at 2K and downscales), does not support 4K (you get the original rate-limit error), and does not support the 4:5 and 5:4 aspect ratios (no fallback happens).
| Request | What the changelog says happens |
|---|---|
| allow_fallback_model true, capacity hit | Runs on Seedream 5.0 lite, billed at Seedream's cost |
| 1K output | Generated at 2K and downscaled |
| 4K output | Returns the original rate limit error |
| 4:5 or 5:4 aspect ratio | Does not fall back |
What does Sume do when a provider is busy?
Sume's image docs list allow_fallbacks among routing fields, then say Sume publishes a single sume endpoint per model in v1, so only and order accept only "sume", and ignore, sort, and allow_fallbacks are accepted and have no effect. Any other slug returns 400 provider_not_available.
Capacity is handled by admission, not by swapping models. Valid paid jobs can be accepted as queued while the workspace has queue capacity, and a full queue returns 429 queue_full. A provider outage can return 503 provider_capacity_exceeded. In every case the model id you submitted is the model that runs.
When is a silent model swap a problem?
If you compare outputs across a batch, a fallback changes the model mid-run and the cost line with it. The Replicate changelog is explicit that billing follows the fallback, which is honest, but your quality review still needs to know which model rendered each frame.
On Sume the swap cannot happen behind your back. If you want a backup model, you make that decision in your own code, where you can log it.
How do you build an explicit fallback on Sume?
Catch the failure, then submit the same prompt to a second model id you chose in advance, with its own idempotency key. A short plan:
- Pick the primary and backup ids from the image catalog, and check each one's supported sizes and aspect ratios first, since a backup that rejects your ratio is no backup.
- Use
Idempotency-Key: <run>-primaryand<run>-backupso retries never double-submit. - Retry the primary on
429 queue_fullor503with backoff before you switch; only a terminal failure justifies the backup. - Record the model id that actually produced each image next to the job id.
What does Sume not do?
It does not offer a switch that reroutes to a cheaper or different model on capacity. It does not publish a per-request fallback list the way OpenRouter does for text. If you want automatic routing, sume/auto lets Sume pick for you, with the trade-off that you do not choose the model.
What does this look like in a cost review?
A fallback makes cost reviews harder in one way and easier in another. Easier: the Replicate changelog says the bill follows the fallback model, so you pay for what ran. Harder: the same prompt can have two prices in one batch, and the cheaper image may not match the look of the others, especially at 1K where Seedream 5.0 lite generates at 2K and downscales.
If you do use it, record the model that ran on every output, and exclude fallback images from any side-by-side test of the primary model.
Sources
Related posts
More in Comparisons
- Replicate predictions source=web filter vs Sume jobs list
Replicate lets you list only web-created predictions, limited to 14 days. Sume's GET /v1/jobs lists only jobs your key's member created. How to scope a list.
- Runway Model Router credit ceiling vs Sume max_spend_usd
Runway's per-modality credit ceiling removes costly models before ranking and errors if none fit. Sume's max_spend_usd caps one call. How the two caps differ.
- Runway Model Router dryRun: preview the pick, and what Sume Auto shows
Runway's Model Router has a dryRun preview, per-modality credit ceilings and cost, latency or quality goals. Sume's sume/auto never says which family ran.
- Runway Seedance 2.0 4K at 150 credits a second: cost
Runway bills Seedance 2.0 4K at 150 credits a second, $1.50 at its $0.01 credit. What a clip costs against 1080p, and what Sume's seedance-2 offers instead.
Written by Sume