Capacity fallbacks: Replicate's model swap vs pinning on Sume

Replicate lists a model falling back to another at capacity. On Sume you pin a model or send sume/auto, and a capacity error is retried with the same key.

4 min readSume
All posts

On Sume you choose the fallback behavior up front: pin a model and handle a capacity error yourself, or send sume/auto and let Sume pick. The Replicate changelog lists a March 2026 change in which Nano Banana Pro can fall back to Seedream 5.0 lite at capacity, so on that platform the swap can happen for you under a single model name.

The difference matters when output has to match across a batch or when you are billed per model.

What the changelog lists

Entries as listed on the page.

Replicate changelog (read 2026-10-03)
ItemWhat the page says
Nano Banana ProCan fall back to Seedream 5.0 lite at capacity (Mar 2026)
PredictionsDeadlines are supported
MCP serverIn the official registry (Feb 2026)

Sume's two modes

With a pinned model, such as seedance-2.5, the request runs on that model. With model: "sume/auto", Sume picks the family and the poll response reports sume/auto. Sume does not disclose which family served the request, and the docs say you should not infer it from any trait of the output.

Resolution is a pure function of the normalized request and the catalog version, so an idempotent replay prices and routes the same way.

Choosing

The right choice depends on what must stay constant.

Pin or Auto on Sume, from the video generation docs (read 2026-10-03)
You needUseTrade-off
The same model across a campaignA pinned catalog idYou handle capacity errors
The job done, any suitable familysume/autoYou cannot tell which family ran
A repeatable priceEither, with dry_run or generation_admission_previewPreview first

What a capacity error looks like on Sume

Two errors look similar and mean different things. 429 queue_full means your workspace has no accepted generation capacity left: wait for running jobs to finish or cancel queued ones, then retry with the same idempotency key. 503 provider_capacity_exceeded means Sume's provider dispatch queue is full: retry later with the same key.

Neither silently changes the model. A pinned request stays pinned, and a retry with the same Idempotency-Key returns the original job if one was created, rather than billing a second one.

A retry wrapper

The loop below retries 429 and 503 responses (which cover both capacity errors but also other causes, so read the error code before relying on it), honors retry-after when present, and keeps the same key on every attempt. It does not switch models; that decision belongs to you.

async function submit(body, key, tries = 5) {
  for (let i = 0; i < tries; i++) {
    const res = await fetch("https://api.sume.com/v1/videos", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.SUME_API_KEY}`,
        "Content-Type": "application/json",
        "Idempotency-Key": key,
      },
      body: JSON.stringify(body),
    });
    if (res.status !== 429 && res.status !== 503) return res;
    const wait = Number(res.headers.get("retry-after")) || 2 ** i * 5;
    await new Promise((r) => setTimeout(r, wait * 1000));
  }
  throw new Error("capacity did not clear");
}

Billing follows the model

Sume bills list price times 1.25 on every video model and reserves the estimate when a job is accepted. A failed job releases or refunds the reservation where applicable, and a successful one is captured. Read the estimate before a large batch, and check the live catalog with GET /v1/videos/models rather than assuming one model's limits apply to another.

Limits differ per model. For example the docs list seedance-2.5 at 4 to 30 seconds and every other catalog model not named otherwise as capped at 15 seconds, so a fallback you design yourself should check duration against the target model before submitting.

When you want a fallback chain

If you want an explicit chain, such as a pinned model first and Auto second, write it yourself and use a different key for the second hop, because a new model is a new operation. Record both job ids, and cancel the first one if it is still queued, since cancellation succeeds only before generation starts.

A plain sume/auto request is simpler when nothing downstream depends on a specific model.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume