Replicate allow_fallback_model: Nano Banana Pro vs Sume

Replicate can fall back from Nano Banana Pro to Seedream 5.0 lite and bills the fallback. Sume's allow_fallbacks is accepted but has no effect. What to do.

5 min readSume
All posts

Replicate's Nano Banana Pro can fall back to ByteDance Seedream 5.0 lite when Google's API is at capacity, but only if you set allow_fallback_model to true, and you are charged the fallback model's cost. Sume has no equivalent: its image API accepts provider.allow_fallbacks, but with one sume endpoint per model the flag has no effect, so a capacity problem shows up as a queued job or an error, not a different model.

Replicate details come from its changelog entry of March 2, 2026, read on 2026-10-03. Sume details come from the image generation docs, errors and credits, and generation admission.

What does Replicate's fallback actually do?

Per the changelog, setting allow_fallback_model to true lets Nano Banana Pro route to Seedream 5.0 lite when Google's API reaches capacity. The entry states: "You're charged the cost of the fallback model, not Nano Banana Pro."

It also lists limits where the fallback cannot help: Seedream 5.0 lite does not support 1K (it generates at 2K and downscales), does not support 4K (you get the original rate-limit error), and does not support the 4:5 and 5:4 aspect ratios (no fallback happens).

Replicate fallback rules, Nano Banana Pro (read 2026-10-03)
RequestWhat the changelog says happens
allow_fallback_model true, capacity hitRuns on Seedream 5.0 lite, billed at Seedream's cost
1K outputGenerated at 2K and downscaled
4K outputReturns the original rate limit error
4:5 or 5:4 aspect ratioDoes not fall back

What does Sume do when a provider is busy?

Sume's image docs list allow_fallbacks among routing fields, then say Sume publishes a single sume endpoint per model in v1, so only and order accept only "sume", and ignore, sort, and allow_fallbacks are accepted and have no effect. Any other slug returns 400 provider_not_available.

Capacity is handled by admission, not by swapping models. Valid paid jobs can be accepted as queued while the workspace has queue capacity, and a full queue returns 429 queue_full. A provider outage can return 503 provider_capacity_exceeded. In every case the model id you submitted is the model that runs.

When is a silent model swap a problem?

If you compare outputs across a batch, a fallback changes the model mid-run and the cost line with it. The Replicate changelog is explicit that billing follows the fallback, which is honest, but your quality review still needs to know which model rendered each frame.

On Sume the swap cannot happen behind your back. If you want a backup model, you make that decision in your own code, where you can log it.

How do you build an explicit fallback on Sume?

Catch the failure, then submit the same prompt to a second model id you chose in advance, with its own idempotency key. A short plan:

  • Pick the primary and backup ids from the image catalog, and check each one's supported sizes and aspect ratios first, since a backup that rejects your ratio is no backup.
  • Use Idempotency-Key: <run>-primary and <run>-backup so retries never double-submit.
  • Retry the primary on 429 queue_full or 503 with backoff before you switch; only a terminal failure justifies the backup.
  • Record the model id that actually produced each image next to the job id.

What does Sume not do?

It does not offer a switch that reroutes to a cheaper or different model on capacity. It does not publish a per-request fallback list the way OpenRouter does for text. If you want automatic routing, sume/auto lets Sume pick for you, with the trade-off that you do not choose the model.

What does this look like in a cost review?

A fallback makes cost reviews harder in one way and easier in another. Easier: the Replicate changelog says the bill follows the fallback model, so you pay for what ran. Harder: the same prompt can have two prices in one batch, and the cheaper image may not match the look of the others, especially at 1K where Seedream 5.0 lite generates at 2K and downscales.

If you do use it, record the model that ran on every output, and exclude fallback images from any side-by-side test of the primary model.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume