Synthesia Syren writes video as code: what Sume has instead
Syren is Synthesia's prompt-to-video agent in early access. Sume has no equivalent; it has a script-to-avatar API and a timeline. The gap, mapped.

Syren is Synthesia's new prompt-driven video tool, and Sume does not have an equivalent. Sume has no agent that learns your brand from past videos and writes a whole motion-graphics video from a chat prompt. What Sume ships is narrower and more deterministic: a script-to-avatar API, first-frame previews, and a timeline that joins clips to one audio track.
Synthesia's own page describes Syren this way: "You describe the video you want, and Syren writes it as code and renders it live in front of you" (read 2026-10-11). It learns a brand kit from existing videos, takes prompts, documents, images or video links, and produces motion graphics, 3D, voiceover, music and avatars. Everything about Syren below comes from that page and its help article, both read 2026-10-11; everything about Sume comes from the linked Sume docs.
What Syren is, from Synthesia's own pages
The help article lists the steps: pick Wide 16:9 or Tall 9:16, describe the video, optionally attach documents, images, videos or PowerPoint files, press Send, watch a live preview, refine through chat, then generate at a chosen resolution. Videos can be up to 3 minutes long. Free videos on the Basic plan carry a watermark. Interactive avatars and live conversations are not available inside Syren.
The product page adds that Syren includes Synthesia's avatars and voices, but your own Personal Avatar or cloned voice requires building the video in Synthesia's regular editor instead.
What Sume offers for the same job
Sume's avatar route turns a ready avatar and a script into a talking video of 4 to 60 seconds. You can send one script, or ordered video_inputs with spoken scenes and silence beats, plus an optional scene prompt or photo and an optional product image (Generate avatar video). Before paying for the full render you can approve first-frame stills (Avatar video previews). To go past 60 seconds, you split the script and join the clips on Timeline 1.0, which takes one audio spine and up to 200 ordered video slots.
Nothing in those pages writes the video for you. You decide the script, the scenes and the cuts; Sume renders them. Avatar 1.0 is also English-only, so a Syren-style multilingual brand video is not something to expect from the avatar route.
Two more differences matter for planning. First, Sume prices avatar video per second before you submit, so a script's cost is known in advance: $0.184, $0.245 or $0.55 per second on the standard, plus and max tiers, with a small per-second surcharge when you add a product image. Syren's page quotes credits instead: a two-minute video with up to two rounds of edits typically uses 1,000 to 2,000 credits, and the pages I read do not convert credits to dollars. Second, Sume's output is a plain MP4 at 720p from the avatar job, and any design layer around it, such as cards or captions, is something you add with the caption option or a timeline.
Side by side
The table compares only what each vendor documents in the pages named in its caption.
| Question | Syren | Sume |
|---|---|---|
| How do you start? | Chat prompt, optional files | Script or video_inputs in an API request |
| Learns a brand kit from your videos? | Yes, per Synthesia's page | Not in the docs I read |
| Longest video | 3 minutes | 60 seconds per avatar job; Timeline joins to 1,800 s |
| Your own face as the presenter | Outside Syren, in Synthesia's regular editor | Avatar from a photo, then reuse its handle |
| Live, interactive avatar | Not available in Syren | Not offered; clips only |
| Approve before the full render | Live preview while chatting | First-frame previews, then generate-video |
How to choose this week
If you want a designed, motion-graphic explainer from a prompt and a pile of brand footage, Syren is the product built for that, and it is in early access with launch credits. If you want a repeatable pipeline where a script in a database becomes a 4 to 60 second English talking clip with a price known before you submit, use Sume's avatar route and a timeline for anything longer.
A practical split: let people prototype a brand look in Syren if they have access, then script the production runs through the API so each clip is reproducible with an idempotency key. For a decision about live avatars instead of clips, read the live versus rendered comparison.
One caution on timing: Syren is early access as of the pages read on 2026-10-11, and Synthesia's launch credits have an end date, so any workflow you build on it this week should be treated as a trial. Sume's avatar endpoints are generally available, which matters if a finance or legal reviewer needs a stable route to cite.
Sources
Related posts
More in Comparisons
- Syren mixes the first 24 audio tracks; Sume's timeline takes 20 parts
Syren mixes only the first 24 audio tracks in an export. Sume's timeline takes one spine in up to 20 gapless parts plus one bed. What fits where.
- Syren can't use your own avatar or cloned voice; Sume has a handle
Synthesia says own avatars and cloned voices need its regular editor, not Syren. Sume builds an avatar from a photo; voice cloning is app-only.
- Syren's 3-minute cap and 30-second avatar scenes vs Sume's 60 s
Syren runs up to 3 minutes but avatar scenes over 30 s can time out; Sume takes 4-60 s per avatar job. How to plan a 3-minute presenter video.
- WaveSpeed levels come from one top-up; Sume concurrency is plan-only
WaveSpeed sets Bronze to Ultra by a single top-up amount. On Sume a top-up never raises concurrency; the plan does (1, 4, 8, 20). What that means for budgets.
Written by Sume