AI model router: what it does and what you give up
An AI model router picks which model serves each request, so your code sends one id. How that differs from pinning a model, and the control you trade.

An AI model router is a layer that chooses which model serves each request. Your code sends one model id, and the router decides which model actually runs. You get a simpler integration; you give up knowing, and controlling, which model produced a given result.
Sume has a router for video and image generation, sume/auto, so it serves as the worked example. Its facts come from the Video Generation, Video Router and Image API docs, read on 2026-09-28.
How is a model router different from picking a model?
The difference is who chooses and whether you find out. A model aggregator gives you many models behind one key but still makes you name one; see AI model aggregator. Watch the names, too: despite its name, Sume's Video Router is an explicit model catalog where you pick the model id, and routing only happens when you send sume/auto.
| Aspect | Pinned model | Model router |
|---|---|---|
| Who picks the model | You | The router |
| What you send on Sume | A catalog id, e.g. seedance-2.5 | "model": "sume/auto" |
| What the response reports | The id you sent | sume/auto; the family that ran is never disclosed |
| Listed in the model catalog | Yes | No: sume/auto is not listed in GET /v1/images/models |
How does a model router API call look?
It looks like any other call, with the router's id in model. This is the POST /v1/videos body from the Sume docs:
- Replays are stable. The docs say resolution is a pure function of the normalized request and the catalog version, so an idempotent replay prices and routes identically.
- It is opaque. The poll response reports
"model": "sume/auto", and the docs say not to build on any observable trait of the output to infer the family. - It has its own limits. Auto create controls default to 720p and 8 seconds, with 3–10 second clips at 16:9 or 9:16.
{
"model": "sume/auto",
"prompt": "A vertical UGC-style product clip on a desk, natural light",
"aspect_ratio": "9:16",
"duration": 5
}What do you give up with a model router?
- Knowing which model ran. With
sume/auto,job.modelstayssume/auto, so you can't log or report the family. - A model's full range. A pinned model can offer more than the router's defaults:
seedance-2.5accepts 4–30 seconds at 480p, 720p or 1080p, against Auto's 3–10 second clips. - A fixed family. Routing follows the request and the catalog version, not a family you chose, so pin a model when a result must stay tied to one family.
Should I use a router or pin a model?
Use the router when any capable model will do and you would rather not track the catalog yourself. Pin a model when you need a specific length or resolution, or need to say which model made a file. An OpenRouter-compatible video API works through that choice for Sume's video endpoint.
Sources
Related posts
More in Developers
- Why an API returns 404 Not Found, and how to fix it
An API returns 404 when no route matches your path or method, or when the id doesn't exist for your credentials. How to tell them apart and fix each.
- API vs SDK: what's the difference, and do you need one?
An API is the contract: routes, fields and errors. An SDK is a library in one language that calls that API for you. What an SDK adds, and when to skip it.
- Arcads API: how it works, from credentials to video
Arcads has a public API: Basic auth with a client ID and secret, then a brand, a folder and a script, one generate call, and a poll for the video URL.
- Asynchronous request-reply pattern: how it works
In the asynchronous request-reply pattern, the server accepts work with a 202 and a status URL, and the client polls or takes a callback until it's done.
Written by Sume