Pin the model id in an ad test: sume/auto follows the catalog

sume/auto is a pure function of the request plus the catalog version, so two ad arms made weeks apart can land on different models. Pin an explicit id in tests.

5 min readSume
All posts

For an ad A/B test, send an explicit model id such as seedance-2 on every arm instead of sume/auto. The Sume /v1/videos contract defines sume/auto as a pure function of the normalized request plus the catalog version. That is good for idempotent replays, but it also means the target can change when the catalog version changes, so two arms generated weeks apart may not share a model. Your test would then compare model changes, not hooks.

What sume/auto does today

The contract says there is no load balancing and no A/B inside auto, so a replay gets the same route and the same price. The default target for text, first-frame or end-frame image, and reference generation is gemini-omni-flash-1.1. Auto validates against that model's envelope and fails closed when a request goes outside it, instead of silently routing elsewhere.

sume/auto envelope from the /v1/videos contract, read 2026-10-07
PropertyValue under sume/auto
Duration3 to 10 seconds, default 8
Resolutions360p, 720p, 1080p, 4K; default 720p
Aspect ratio16:9 or 9:16
Native audioAlways on; generate_audio false fails
Resolved model in responsesHidden; job.model echoes sume/auto
Billing keyThe resolved family, since auto has no price of its own

Why the hidden model hurts a test

The contract says the resolved family must not appear in public response bodies, headers, errors or catalog routing metadata. For production use this keeps your integration stable. For a test it removes the one fact you need to prove the arms are comparable. If two variants both say model sume/auto, you cannot tell from the response which family served each.

The legacy Video 1.0 compatibility route records request.resolved_model internally, but that is not something to build a test around. An explicit catalog id is echoed back verbatim, so the poll response itself documents which model made the clip.

When auto is fine

Auto is the right choice when you want Sume to pick for a one-off clip, or when an agent calls the MCP generate_video tool without a model. It is also fine for finals when the test is over and you only want a good result inside Omni's envelope.

Pin an id when any of these are true:

  • You compare hooks, endings or captions and need the model to be a constant.
  • You need 1:1, 4:3 or 3:4 output, which Omni's 16:9 and 9:16 envelope does not offer.
  • You need a duration outside 3 to 10 seconds.
  • You need generate_audio false, which auto rejects.

A pinned arm body

Here is one arm of a hook test. Change only the prompt between arms, and keep model, duration and aspect_ratio fixed. Check the id against GET /v1/videos/models before you start so the duration and ratio are advertised.

If you later move the whole test to another model, record that change in your sheet and rerun every arm. A model switch in the middle of a test is a second variable, and no amount of statistics fixes that.

curl -sS -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: arm-b-hook-v1" \
  -d '{"model":"seedance-2","prompt":"Hook B: a hand drops a coin into a jar","duration":6,"aspect_ratio":"9:16","resolution":"720p"}'

A pre-flight check you can script

Before the first submit, read GET /v1/videos/models and confirm that your pinned id lists the duration, aspect ratio and resolution you plan to send. The catalog is a projection of the Video Router capability records, so it cannot disagree with what validation enforces. If a value is missing the API gives a 400 for that field, and the same failure on a sume/auto arm would name sume/auto and hide which family it hit.

Record the catalog read in your test log with its date. If the model list or its supported values change between your first and last arm, you will know that the environment moved, and you can rerun the early arms instead of guessing.

The cost of being wrong

A pinned id does not make results identical, because Sume's video route does not accept a seed: size and seed are 400 unsupported_parameter, and each model reports seed false. Two runs of the same prompt on the same model will still differ. That is why a pinned model matters more, not less. The only variation you want is the prompt and the natural spread of the model, not a second model on top of it.

If your budget allows it, run each arm twice and keep the better take, using the same prompt. That reduces the chance that one unlucky render decides a hook.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume