Synthesia test videos are free; Sume previews first frames instead
Synthesia's API test flag renders free watermarked drafts. Sume has no test flag; avatar video previews approve first frames before the full render.

Synthesia's Create a video endpoint has a test boolean, default false, and its reference says test videos are free and not counted toward your quota. Sume has no equivalent flag on the talking-video route. What it offers instead is a first-frame stage: an avatar video preview renders stills, you approve them, and only then do you call generate-video on the preview id. The two protect your budget in different ways.
Synthesia facts below come from its create-video reference and pricing page, read on 2026-10-04. Sume facts come from Generate avatar video and the preview page.
What a Synthesia test video gives you
A test render goes through the same pipeline as a real one, so you see motion, voice and timing. The reference notes that test videos carry a watermark overlay. That makes the flag a good way to check script wiring and layout in CI without spending quota, but the artifact is not something you can ship.
| Question | Synthesia | Sume |
|---|---|---|
| How you ask for a draft | test: true on the create call | POST /v1/avatar-video-previews |
| What you get back | A full video with a watermark overlay | First-frame stills (preview_image_url, scene_previews[]) |
| Is motion or voice included | Yes, per the reference | No, stills only |
| Cost of the draft | Free, not counted toward quota | Check current pricing for the preview job before relying on it |
| Moving from draft to final | New create call with test false | generate-video on the preview id, same first frame reused |
What the Sume preview stage gives you
A preview takes the same body as the talking-video route: exactly one of script or video_inputs, plus optional product_image, scene, quality, aspect_ratio, title and captions. The response carries an avatar_video_preview_id, and you poll the job like any other generation.
Because the still is reused, the final render keeps the framing you approved. Preview stills are tier-independent, so you can approve once and choose standard, plus or max only when you call generate-video. Structural changes such as a new script need a new preview.
Which draft fits which risk
Use a watermarked full render when the risk is timing, pacing or voice. Use stills when the risk is composition: wrong scene, wrong product placement, a crop that hides the face. Sume's stage cannot tell you whether a 40-second script feels slow, so read the script aloud and keep it inside the 4 to 60 second window the docs describe.
- Run a preview for every new scene prompt or product image.
- Skip it for repeat scripts on a scene you already approved.
- Store the preview id with your own record so a reviewer can approve later without re-rendering.
A minimal Sume call
The order is preview, poll, read the resource, then generate. Read resource_status for readiness and job_status for polling, as the docs recommend over the legacy status field.
curl -X POST https://api.sume.com/v1/avatar-video-previews \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: preview-launch-001" \
-d '{"avatar_handle":"product_host","script":"Meet the new travel mug.","aspect_ratio":"9:16"}'Bottom line
If your workflow depends on a free full-motion draft, Synthesia's test flag does that and Sume does not. If you need an approval gate before paying for the final render, Sume's preview stage is built for it.
Sources
Related posts
More in Comparisons
- Tavus lists 1080p conversational video; Sume avatar clips are 720p
Tavus's pricing page lists 1080p and 24 kHz audio on all plans; Griffin-Lite renders 720p. Sume avatar videos are 720p today. What that means for your delivery.
- Is there a Tavus Griffin-Lite API? Not yet, and what to use
Tavus says Griffin is not available to customers, only to select trusted testers. Sume's avatar video API ships today with 4-60 s clips and five ratios.
- Tavus Starter: 3 concurrent streams vs Sume bulk concurrency 1 to 16
Tavus Starter caps live conversations at 3 concurrent streams. Sume bulk runs queue up to 100 renders and keep 1 to 16 in flight. Different jobs.
- TTS language counts in October 2026: Voxtral 9, MAI 23, ElevenLabs 90+
How many languages each TTS vendor states: Mistral Voxtral, Microsoft MAI-Voice-2.1, OpenAI and ElevenLabs, from vendor pages, plus Sume's language field.
Written by Sume