23-second AI video with audio references on Wan 3.0 and Seedance
Audio references are accepted by Wan 3.0 (up to 5 clips, 15 s total) and Seedance 2.5. A 23 s clip price on Wan and the rules for audio inputs.

For a 23 second clip that follows an audio reference, use Wan 3.0 or Seedance 2.5 on Sume. Wan accepts up to 5 audio clips with a combined length of 15 s. Seedance 2.5 lists audio references among its capabilities.
Which models accept 23 seconds
Higgsfield Genjutsu and H3 Max Recast take images and a source video instead of audio references.
| Model id | Duration range | Billed for 23 s |
|---|---|---|
| wan-3.0 | 2-30 s | 480p $1.44, 720p $2.88, 1080p $5.75 |
| seedance-2.5 | 4-30 s | token priced, read pricing_skus |
| higgsfield-genjutsu | 4-30 s (source length) | read pricing_skus |
| h3-max-recast | 5-30 s (source length) | 768p $8.63, 1080p $12.94 |
Audio references need a visual anchor
On the legacy video wire, an audio reference needs an image or video reference alongside it; audio cannot be the only input. MiniMax H3 follows the same rule: audio cannot be the only reference.
The audio total is capped at 15 s on Wan, so a 23 s output will outlast its audio reference. Plan for that: trim to what you need, or add a second beat in the prompt.
What a 23 second request costs
Sume bills the provider list price times 1.25, rounded up to the next cent. The amount is reserved when you submit and refunded if the job fails, so a rejected or failed clip does not cost you. The /v1/videos/models descriptor lists the billable rate in pricing_skus, which is the field to trust over any number in a blog post.
A 23 s Wan clip bills $1.44 at 480p, $2.88 at 720p and $5.75 at 1080p.
Check 23 seconds against the live catalog
Duration support lives in the catalog, not in this page. Each entry on GET /v1/videos/models carries supported_durations, which is every whole second from the model minimum to its maximum, so one membership test answers the question. This script lists every model that accepts 23 s, with its resolutions and billable rates.
Set SUME_API_KEY first. See the Video Generation docs for the request and polling lifecycle.
import os
import requests
N = 23
resp = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
resp.raise_for_status()
for m in resp.json()["data"]:
if N in m["supported_durations"]:
print(m["id"], m["supported_resolutions"], m["pricing_skus"])Sources
Related posts
More in Developers
- 502 from the Sume Images API: read code, next_action and status_url
A 502 on POST /v1/images is a failed job, not an outage. Parse error.code, retryable and next_action, then fetch status_url for the job. Python handler.
- MiniMax H3 on Sume: "720p is not supported; use 768p" explained
MiniMax H3 and H3 Max render natively at 480p and 768p. Send resolution 720p and Sume answers 400. Why, what to send, and what it costs per second.
- A/B test two reference sets on Seedance 2.5: six takes for about $16
Which reference images work better on Seedance 2.5? Run two sets, three takes each, 10 s at 480p on Sume for about $16, and score blind.
- Idempotency-Key for Agent Completions: reuse the model's tool call id
Retries after a timeout must not start a second paid Sume run. Derive Idempotency-Key from the tool call id your model returned, such as the OpenAI call_id.
Written by Sume