AI model router: what it does and what you give up

An AI model router picks which model serves each request, so your code sends one id. How that differs from pinning a model, and the control you trade.

5 min readSume
All posts

An AI model router is a layer that chooses which model serves each request. Your code sends one model id, and the router decides which model actually runs. You get a simpler integration; you give up knowing, and controlling, which model produced a given result.

Sume has a router for video and image generation, sume/auto, so it serves as the worked example. Its facts come from the Video Generation, Video Router and Image API docs, read on 2026-09-28.

How is a model router different from picking a model?

The difference is who chooses and whether you find out. A model aggregator gives you many models behind one key but still makes you name one; see AI model aggregator. Watch the names, too: despite its name, Sume's Video Router is an explicit model catalog where you pick the model id, and routing only happens when you send sume/auto.

Sume behavior from Video Generation, Video Router and Image API, read 2026-09-28.
AspectPinned modelModel router
Who picks the modelYouThe router
What you send on SumeA catalog id, e.g. seedance-2.5"model": "sume/auto"
What the response reportsThe id you sentsume/auto; the family that ran is never disclosed
Listed in the model catalogYesNo: sume/auto is not listed in GET /v1/images/models

How does a model router API call look?

It looks like any other call, with the router's id in model. This is the POST /v1/videos body from the Sume docs:

  • Replays are stable. The docs say resolution is a pure function of the normalized request and the catalog version, so an idempotent replay prices and routes identically.
  • It is opaque. The poll response reports "model": "sume/auto", and the docs say not to build on any observable trait of the output to infer the family.
  • It has its own limits. Auto create controls default to 720p and 8 seconds, with 3–10 second clips at 16:9 or 9:16.
{
  "model": "sume/auto",
  "prompt": "A vertical UGC-style product clip on a desk, natural light",
  "aspect_ratio": "9:16",
  "duration": 5
}

What do you give up with a model router?

  • Knowing which model ran. With sume/auto, job.model stays sume/auto, so you can't log or report the family.
  • A model's full range. A pinned model can offer more than the router's defaults: seedance-2.5 accepts 4–30 seconds at 480p, 720p or 1080p, against Auto's 3–10 second clips.
  • A fixed family. Routing follows the request and the catalog version, not a family you chose, so pin a model when a result must stay tied to one family.

Should I use a router or pin a model?

Use the router when any capable model will do and you would rather not track the catalog yourself. Pin a model when you need a specific length or resolution, or need to say which model made a file. An OpenRouter-compatible video API works through that choice for Sume's video endpoint.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume