Vidu API: S2-Avatar live voice vs Sume avatar video jobs
Vidu S2-Avatar is a real-time voice model. Sume has no live session: you submit a script and a scene photo to an avatar job and fetch the finished video.
If you searched "Vidu API" after the S2-Avatar launch: Vidu describes a real-time interactive model, and Sume does not ship a live session. Sume's avatar endpoint renders a script into a finished talking video, which you fetch when the job ends.
Vidu facts are from its platform update log; Sume facts are from Generate avatar video and Webhooks, all read 2026-09-30.
What does Vidu say S2-Avatar does?
The September 15, 2026 entry lists "Vidu S2-Avatar: Real-Time Interactive Model". It supports real-time voice interaction and complex motion control, and reference images can be used for product interaction, outfit changes and background replacement. The page says nothing more about request fields, so check Vidu's own docs before you plan around it.
What does Sume do instead?
Sume's docs say avatar videos "turn a ready avatar into a script-driven talking video" through POST /v1/avatar-1.0/talking-video. You send an avatar_handle plus exactly one of script or video_inputs. The result is a file, not a stream, so there is no way to speak to it mid-render.
For a background, set scene to { "type": "photo", "image_url": "https://..." } for a photo reference, or { "type": "prompt", "prompt": "..." } for scene direction. The image must be a fetchable public HTTPS URL.
How do the two compare on the points Vidu lists?
| Need | Vidu S2-Avatar (vendor page) | Sume avatar video (docs) |
|---|---|---|
| Live voice interaction | Listed | Not offered; script in, video out |
| Background change | Reference images | scene photo or prompt |
| Product in frame | Reference images | Optional product_image |
| Delivery | Real-time session | Async job, status poll or webhook |
How do I get the finished video?
Poll the job or use a webhook. Sume sends terminal job events only: job.completed, job.failed and job.canceled, with no progress or partial deliveries. Scripts must land at 4-60 seconds of estimated duration, so split longer ones into several jobs.
When should I pick a live model instead?
If a viewer must talk to the avatar and hear an answer immediately, you need a live session product, and Sume's rendered jobs will not cover that. If you can script the lines ahead of time, a rendered job works; see real-time AI avatar vs video avatar API for the longer comparison.
Sources
Related posts
More in Models
- Vidu S2-Editing: live stream edits vs a Sume Omni clip edit
Vidu S2-Editing edits an incoming video stream in real time. Sume edits one finished clip with a prompt through Gemini Omni video_to_video, as an async job.
- Wan 3.0 on Runway, or the wan-3.0 id as an API job on Sume
Runway's changelog lists Wan 3.0 in tool mode and workflows on paid plans. On Sume, wan-3.0 is a model id you call from code and poll as a job.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
Written by Sume