Griffin-Lite inputs vs a Sume talking-video request, side by side
Tavus describes a reference image plus streaming audio and controls; Sume takes an avatar handle, a script and a tier. The two request shapes compared.
Tavus says Griffin-Lite works from a reference image together with streaming audio and controls. A Sume talking-video request works from a ready avatar handle together with a finished script (or ordered scenes), a quality tier and optional product and scene media. One is an input stream that never ends while a conversation lasts; the other is a document you submit once and read back.
| Input | Tavus Griffin-Lite | Sume Avatar 1.0 talking video |
|---|---|---|
| Identity | Reference image | avatar_handle created from prompt, props or photo |
| Speech | Streaming audio, packets as small as 10 ms | script or video_inputs text, or silence beats |
| Direction | Controls | scene prompt or photo, product_image, aspect_ratio, quality |
| Completion | Continuous | Job: queued, processing, completed, failed, canceled |
| Access | Select trusted testers | Public API, Idempotency-Key |
Identity: image or handle
For Sume the reference image is used once, at avatar creation. You send input.type: "photo" with a fetchable public HTTPS image_url, and the avatar becomes a named resource. Later requests carry only the handle. Sume rejects localhost, private-network and non-HTTPS URLs and non-image responses before it submits the generation.
That is a different lifecycle from sending an image into a live model each session. The handle makes the same face reusable across many jobs and across routes.
Speech: stream or script
In a conversation the audio arrives as the other person speaks. In a Sume request the words are known in advance and Sume plans the video: it estimates duration and requires 4 to 60 seconds, splits spoken text into clips of 4 to 12 seconds, and rejects a script that is too long with a message to split it.
Silence is explicit. A scene with voice.type: "silence" needs a duration and no script.
Direction: controls or fields
Griffin's page mentions controls without a published list in what we read. Sume's controls are documented fields: quality of standard, plus or max, aspect_ratio of 1:1, 3:4, 9:16, 4:3 or 16:9, a 720p resolution, an optional product image, and a prompt or photo scene. Plan around the documented ones, and check the live OpenAPI reference for the full schema.
What to take from the comparison
The point is not that one shape is better. A stream suits a conversation, a document suits a production. If you are weighing the two, the useful question is how much of your content you can know in advance. The more you can write down, the more a request like Sume's saves you: it is reviewable, repeatable and billed by the second at a published rate.
Whatever you choose, keep the identity inputs consistent. Use one reference image per person, with their permission, and do not mix photos of different people under one name.
Sources
Related posts
More in Comparisons
- Haiku 5.5 vs Sonnet 5.5 token prices for an agent calling Sume
Haiku 5.5 costs $0.10 per million input tokens and Sonnet 5.5 costs $2. What the gap does and does not change for a Sume agent's spend cap.
- HappyHorse 1.1 is 12th at 1,042 Elo: what ranks above it on Sume
HappyHorse-1.1 is 12th on the AA text-to-video board at $10.80 a minute. Sume does not list it; six Sume models rank above it. Per-minute prices inside.
- HappyScribe $0.20 a minute vs a $0.20 Sume caption job
HappyScribe tops up AI minutes at $0.20 each. Sume burns captions into a clip of up to 60 seconds for $0.20 flat. Where each wins.
- Hedra's developer API is in private beta; the Sume routes open to call
Hedra lists its developer API as a private beta. Sume documents its avatar routes publicly: lip sync, Fabric, talking video, motion control. Routes and limits.
Written by Sume