Griffin generates the whole scene: what a Sume avatar scene is

Tavus says Griffin generates body, background and shadows live. Sume avatar videos take a scene prompt or photo, one shared scene per video. What you control.

4 min readSume
All posts

What Tavus says Griffin produces

Tavus's Griffin page, read 2026-10-03, says the model does full-scene generation: not only the face, but the body, the background and the shadows, in real time, along with its own voice and responses. You do not supply a backdrop for a live call, and you cannot art-direct each frame.

Griffin-Lite is open to select trusted testers only, so this is a description of a preview.

What you control in a Sume avatar scene

Sume's avatar video is the opposite: you set the scene up front. The scene field is either { "type": "prompt", "prompt": "..." } for direction in words or { "type": "photo", "image_url": "https://..." } for a photo reference. A product_image is optional, and omitting it gives a productless clip.

Multi-scene plans use video_inputs, and each scene can carry a background of type prompt or image. Current execution expects the backgrounds to resolve to one shared scene and uses one resolved avatar per video, so later scenes are the same room and the same person.

Scene control, from each vendor's page (read 2026-10-03)
QuestionGriffinSume avatar video
Who sets the backgroundThe model, liveYou, by prompt or photo
Can you review it firstNo, it is a callYes, with a first-frame preview
Scene changes mid-videoNot describedOne shared scene per video
Lighting and shadowsGenerated with the sceneFollow the scene prompt or photo

Review before you pay for the render

Because the scene is chosen up front, the preview endpoint makes sense: create first-frame stills, check the framing and background, then generate the video from the preview id. The preview's stills are tier-independent, so you can approve one and pick the render quality at the final step.

Changing the script, the scene or the aspect ratio after approval needs a new preview. Regenerating only refreshes the stills for the same request.

Choosing between them

If the setting matters, such as a real shop floor or a product on a desk, a photo scene gives you a place you picked. If you need a presenter who adapts as a person talks, that is Griffin's category, and Sume does not offer it.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume