Let the API pick the video model: sume/auto for vertical UGC clips

Send model sume/auto to POST /v1/videos and Sume picks the family. The response echoes sume/auto and never names the model. When to pin a model instead.

4 min readSume
All posts

To let the API pick the video model, set model to sume/auto on POST /v1/videos. Sume chooses the family for you, the poll response reports "model": "sume/auto", and it does not disclose which family ran. For a vertical UGC-style clip, add aspect_ratio: "9:16" and a short duration.

Behavior is from the Video generation docs, read 2026-09-29.

What does the request look like?

This is the docs' own example shape. Idempotency-Key makes retries safe: a replay returns the original job.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: auto-ugc-001" \
  -d '{
    "model": "sume/auto",
    "prompt": "A vertical UGC-style product clip on a desk, natural light",
    "aspect_ratio": "9:16",
    "duration": 5
  }'

What is predictable about it?

  • The choice is a pure function of the normalized request and the catalog version, so an idempotent replay routes and prices identically.
  • Billing is reserved on submit at provider list price × 1.25, like every model.
  • You cannot infer the family from the output, and the docs say not to build on any observable trait that hints at it.

When should I pin a model?

Choosing between sume/auto and a pinned model, from Video generation, read 2026-09-29.
You needUse
Any decent clip, no preferencesume/auto
More than 15 secondsseedance-2.5 (4 to 30 s) or wan-3.0 (2 to 30 s)
Audio or video reference inputsA model whose supported_input_references lists them
A fixed look across a whole batchPin one model id

What else does it not do?

It does not change the rest of the contract: size, seed and provider.options are still rejected, and frame_images still takes precedence over input_references. See UGC ad videos via API for the ad-side workflow.

Where do I find what each model supports?

GET /v1/videos/models returns each model's supported_resolutions, supported_aspect_ratios, supported_durations, supported_frame_images, supported_input_references and generate_audio. Check those before pinning. Limits differ: seedance-2.5 accepts 4 to 30 seconds and wan-3.0 2 to 30, while minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, and every other catalog model tops out at 15.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume