Griffin generates the whole scene: what a Sume avatar scene is
Tavus says Griffin generates body, background and shadows live. Sume avatar videos take a scene prompt or photo, one shared scene per video. What you control.
What Tavus says Griffin produces
Tavus's Griffin page, read 2026-10-03, says the model does full-scene generation: not only the face, but the body, the background and the shadows, in real time, along with its own voice and responses. You do not supply a backdrop for a live call, and you cannot art-direct each frame.
Griffin-Lite is open to select trusted testers only, so this is a description of a preview.
What you control in a Sume avatar scene
Sume's avatar video is the opposite: you set the scene up front. The scene field is either { "type": "prompt", "prompt": "..." } for direction in words or { "type": "photo", "image_url": "https://..." } for a photo reference. A product_image is optional, and omitting it gives a productless clip.
Multi-scene plans use video_inputs, and each scene can carry a background of type prompt or image. Current execution expects the backgrounds to resolve to one shared scene and uses one resolved avatar per video, so later scenes are the same room and the same person.
| Question | Griffin | Sume avatar video |
|---|---|---|
| Who sets the background | The model, live | You, by prompt or photo |
| Can you review it first | No, it is a call | Yes, with a first-frame preview |
| Scene changes mid-video | Not described | One shared scene per video |
| Lighting and shadows | Generated with the scene | Follow the scene prompt or photo |
Review before you pay for the render
Because the scene is chosen up front, the preview endpoint makes sense: create first-frame stills, check the framing and background, then generate the video from the preview id. The preview's stills are tier-independent, so you can approve one and pick the render quality at the final step.
Changing the script, the scene or the aspect ratio after approval needs a new preview. Regenerating only refreshes the stills for the same request.
Choosing between them
If the setting matters, such as a real shop floor or a product on a desk, a photo scene gives you a place you picked. If you need a presenter who adapts as a person talks, that is Griffin's category, and Sume does not offer it.
Sources
Related posts
More in Sume Avatar 1.0
- Griffin ranks second on LSE-C: how to judge lip sync yourself
Tavus says LSE-C rewards pronounced mouth movement beyond naturalness. A way to check a Sume avatar clip with timed stills and a word-timed transcript.
- HeyGen lipsync captions are always on; Sume's are opt-in
HeyGen deprecated enable_caption on translation and lipsync and now always returns SRT and VTT. Sume avatar videos burn captions only when you ask for them.
- Holiday avatar ad roster: 3 presenters x 4 scripts, cost by tier
Three reusable avatars and four holiday scripts make 12 clips. Priced on standard, plus and max with a product image; the avatars cost $2.85.
- LemonSlice API: image plus streaming audio vs Sume lip sync
LemonSlice drives a live avatar from an image and streaming audio. For a recorded line, Sume's lip sync takes a still plus an audio_url and returns a file.
Written by Sume