Face swap quality tiers: pick for a 12-second clip
Avatar Face Swap Beta needs quality standard, plus or max and a source clip of about 4 to 15 seconds with audio. A 12-second clip fits; test the tiers.

For a 12-second clip, any of the three tiers is in range: Avatar Face Swap (Beta) wants a source video of about 4 to 15 seconds with usable audio, and quality must be standard, plus or max. The docs do not say how the tiers differ, so choose by rendering the same clip at each tier and comparing.
What the docs fix
The endpoint is POST /v1/models/sume/avatar-face-swap/v1.0/runs. It needs a ready avatar handle and a public HTTPS source video. There is no default for quality on this endpoint.
| Field | Required | Notes |
|---|---|---|
| avatar_handle | Yes | A ready Avatar 1.0 identity |
| video_url | Yes | Public HTTPS video; about 4-15 s with usable audio |
| quality | Yes | standard, plus or max; no omit default |
What the docs do not say
There is no description of what separates standard, plus and max, and no price per tier in the face-swap page. Do not assume that max is better for your clip or that it costs a fixed multiple. Read the live catalog and OpenAPI for pricing, and treat any difference as something to measure.
Run all three on one clip
A 12-second clip fits inside the 4 to 15 second target, so it is a good test unit. Use a separate Idempotency-Key per tier, because they are different requests. Prompts, aspect ratio, duration knobs and provider fields are not supported on this endpoint.
for Q in standard plus max; do
curl -X POST https://api.sume.com/v1/models/sume/avatar-face-swap/v1.0/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: swap-12s-$Q" \
-d "{\"avatar_handle\":\"studio_presenter\",\"video_url\":\"https://example.com/source-12s.mp4\",\"quality\":\"$Q\"}"
doneCompare and decide
Poll GET /v1/jobs/{id}/status and read GET /v1/jobs/{id}/result. Prefer resource_status for readiness and job_status for polling. Watch the clip at full size, check that the audio is intact, and compare the receipt for each tier.
- Keep the source clip, avatar and tier as your only variables.
- Judge faces in motion and at the cuts, not on a single frame.
- Pick the lowest tier that passes your own bar, since you pay for each run.
Source clip checklist
The source must be a fetchable public HTTPS URL. Localhost, private-network, non-HTTPS and signed or private URLs are rejected, and so are provider task URLs. A silent clip is a poor fit, because the Beta targets sources with usable audio.
Sources
Related posts
More in Sume Avatar 1.0
- FTC reviews rule: virtual influencers allowed, fake reviews not
The FTC's reviews rule Q&A says section 465.2 bans fake reviews but does not prohibit virtual influencers. Sume Avatar 1.0 makes spokesperson clips.
- Avatar V multi-look vs Sume scene prompts
HeyGen Avatar V adds multi-look generation. On Sume an avatar keeps one identity, and each video's look comes from scene prompts, product images and quality.
- Italy deepfake offense: 1-5 years, and avatar consent records
A secondary source says Italy's Law 132/2025 took effect Oct 10, 2025 and Art. 612-quater carries 1-5 years for deepfakes. Keep consent records for avatars.
- Split a 3-minute script into 60-second avatar jobs
Sume avatar talking videos accept scripts of an estimated 4-60 seconds. Split a 3-minute script into jobs of that size, then join the audio with timeline audio.
Written by Sume