A/B test two video models: pin the ids, sume/auto hides the family
sume/auto always reports model sume/auto and never discloses the family that ran, so it cannot A/B two models. Pin both ids; a 5 s 480p pair costs $0.76.

Pin the model ids. sume/auto cannot tell you which family produced a clip: the poll response reports "model": "sume/auto", and the Sume docs say Sume never discloses the family that ran and that you should not infer it from traits of the output. A fair A/B of, say, Wan 3.0 against Seedance 2 Mini therefore needs model set to each catalog id, with every other field identical.
What Auto does and does not promise
Auto's selection is a pure function of the normalized request and the catalog version, so an idempotent replay gets the same price and the same route. That is a repeatability guarantee, not a disclosure. The create controls default to 720p and 8 s, accept 3 to 10 s, and offer 16:9 or 9:16. If your test needs a 21:9 frame or a clip longer than 10 s, Auto does not offer it anyway.
A pinned pair, with the price known up front
Hold prompt, resolution, duration and aspect ratio fixed. At 480p, 5 s and 9:16 the two clips below cost $0.76 together.
| Model | Price | Calculation |
|---|---|---|
| wan-3.0 | $0.32 | 0.05 x 5 x 1.25 = 0.3125, rounded up |
| seedance-2-mini | $0.44 | token rate x 1.25, rounded up |
| seedance-2-fast | $0.71 | token rate x 1.25, rounded up |
| seedance-2.5 | $1.35 | token rate x 1.25, rounded up |
The script
This sends the same body to two pinned models. The Idempotency-Key is stable per model, so a retry after a timeout returns the original job instead of a second one. Download from unsigned_urls[0] once the job reports completed.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"}
URL = "https://api.sume.com/v1/videos"
BODY = {"prompt": "A ceramic mug on a desk, slow push-in, natural light",
"resolution": "480p", "duration": 5, "aspect_ratio": "9:16"}
def run(model):
h = {**H, "Idempotency-Key": "ab-test-" + model + "-001"}
r = requests.post(URL, headers=h, json={**BODY, "model": model}, timeout=60)
r.raise_for_status()
poll = r.json()["polling_url"]
while True:
j = requests.get(poll, headers=H, timeout=60).json()
if j["status"] in ("completed", "failed", "cancelled"):
return j
time.sleep(30)
for m in ("wan-3.0", "seedance-2-mini"):
j = run(m)
print(m, j["status"], j.get("usage", {}).get("cost"), j.get("unsigned_urls"))Keep the test honest
- Change one variable. If you change the model and the prompt together you learn nothing about the model.
- Use the same first frame for both if you test image-to-video; frame handling is defined in Video generation.
- Compare at the resolution you will ship. A 480p win does not carry over to 1080p automatically.
- Read
supported_durationsandsupported_aspect_ratiosfromGET /v1/videos/modelsfor each id, since limits differ by model.
When to go back to Auto
After you pick a winner, you can keep that id pinned, or return to sume/auto for traffic where the exact family does not matter. The docs describe Auto as the choice when you do not want to pin a family, and the Video Router as the choice when you do.
Keep the test fair. Send the same prompt, the same start image, the same duration, resolution and aspect ratio to each pinned id, and record the price of each row before you run it. If a model does not accept a setting, such as a duration outside its range, change the test rather than letting one side silently use a different value. Run each pair more than once, since two single generations can differ by chance, and store the job ids so a later review can reopen the exact outputs. Only after that should you decide whether routing through sume/auto is acceptable for production traffic, knowing the response will never tell you which family produced a given clip.
Sources
Related posts
More in Comparisons
- First paid AI avatar plan, October 2026: $18 to $59 vs Sume seconds
Cheapest paid plans at Synthesia, HeyGen, Colossyan and Tavus from their pricing pages, with the seconds of Sume Standard avatar video the same money buys.
- AI avatar news, October 8, 2026: what is worth acting on
Synthesia Sessions and plans, Tavus Griffin-Lite, HeyGen LiveAvatar, Teams and Zoom avatars: dated, with what a buyer can actually do this week.
- Amazon lists 3840x2160, but Sume output tops out at 2160: what to do
Amazon Sponsored Brands video allows 3840x2160. Sume Timeline sides cap at 2160, so a 4K UHD frame is out. Here is the 1920x1080 route and what you give up.
- Amazon no-text lower-right rule vs Walmart contrast: where captions go
Amazon Sponsored Brands bans text in the lower right; Walmart wants 4.5:1 contrast. How to move Sume burned-in captions up with anchor ratios and check a still.
Written by Sume