Griffin-Lite inputs vs a Sume talking-video request, side by side

Tavus describes a reference image plus streaming audio and controls; Sume takes an avatar handle, a script and a tier. The two request shapes compared.

5 min readSume
All posts

Tavus says Griffin-Lite works from a reference image together with streaming audio and controls. A Sume talking-video request works from a ready avatar handle together with a finished script (or ordered scenes), a quality tier and optional product and scene media. One is an input stream that never ends while a conversation lasts; the other is a document you submit once and read back.

Inputs compared, Tavus Griffin page read 2026-10-08, Sume docs as of 2026-10-08
InputTavus Griffin-LiteSume Avatar 1.0 talking video
IdentityReference imageavatar_handle created from prompt, props or photo
SpeechStreaming audio, packets as small as 10 msscript or video_inputs text, or silence beats
DirectionControlsscene prompt or photo, product_image, aspect_ratio, quality
CompletionContinuousJob: queued, processing, completed, failed, canceled
AccessSelect trusted testersPublic API, Idempotency-Key

Identity: image or handle

For Sume the reference image is used once, at avatar creation. You send input.type: "photo" with a fetchable public HTTPS image_url, and the avatar becomes a named resource. Later requests carry only the handle. Sume rejects localhost, private-network and non-HTTPS URLs and non-image responses before it submits the generation.

That is a different lifecycle from sending an image into a live model each session. The handle makes the same face reusable across many jobs and across routes.

Speech: stream or script

In a conversation the audio arrives as the other person speaks. In a Sume request the words are known in advance and Sume plans the video: it estimates duration and requires 4 to 60 seconds, splits spoken text into clips of 4 to 12 seconds, and rejects a script that is too long with a message to split it.

Silence is explicit. A scene with voice.type: "silence" needs a duration and no script.

Direction: controls or fields

Griffin's page mentions controls without a published list in what we read. Sume's controls are documented fields: quality of standard, plus or max, aspect_ratio of 1:1, 3:4, 9:16, 4:3 or 16:9, a 720p resolution, an optional product image, and a prompt or photo scene. Plan around the documented ones, and check the live OpenAPI reference for the full schema.

What to take from the comparison

The point is not that one shape is better. A stream suits a conversation, a document suits a production. If you are weighing the two, the useful question is how much of your content you can know in advance. The more you can write down, the more a request like Sume's saves you: it is reviewable, repeatable and billed by the second at a published rate.

Whatever you choose, keep the identity inputs consistent. Use one reference image per person, with their permission, and do not mix photos of different people under one name.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume