What resolution is Tavus Griffin? 720p, and Sume avatar clips
Tavus says Griffin generates 720p video in 320 ms chunks. Sume avatar video renders at 720p too, as a clip. What the same number means in each product.
The number Tavus gives
Tavus's Griffin page, read 2026-10-03, says the model generates 720p video in 320 ms chunks and reports an average audio-to-video latency of 0.43 seconds on H100 GPUs. It is a full-duplex video-to-video model: it takes audio and video in and produces the face, voice and responses live, in a call.
Griffin-Lite is the research preview, offered to select trusted testers. The same page says it is not available to customers on the Tavus platform for now.
The number Sume gives
Sume's avatar video endpoint renders a finished file. Its docs state that resolution is currently 720p, so the pixel height matches Griffin's stated output. The aspect ratios are 1:1, 3:4, 9:16, 4:3 and 16:9, with 9:16 the default.
The script or multi-scene plan must come to 4 to 60 seconds. Quality is standard, plus (the default) or max, and the tier changes the render path, not the pixel count.
Same height, different product
| Question | Tavus Griffin | Sume avatar video |
|---|---|---|
| Output height | 720p | 720p |
| How it arrives | Live, in 320 ms chunks | One MP4 after a job finishes |
| Who speaks | The model decides in real time | A script you write |
| Length | A call | 4 to 60 seconds per job |
| Access today | Select trusted testers | Public API |
What to do with that
A matching resolution means a Sume clip and a Griffin call can sit in the same page layout without one looking soft next to the other. It says nothing about conversation. Griffin reacts to what you say as you say it; a Sume clip says what you wrote.
If your use is a recorded message on a site, an email or a feed, resolution is not the deciding factor, and 720p at 9:16 or 16:9 covers the common placements. If you need a live answer, Sume has no equivalent today, and the honest plan is to render the scripted parts now and revisit when Griffin is open.
Placement tips
Pick the aspect ratio for the place the clip will play. Sume lists 9:16 for vertical feeds, 16:9 for slides and webinars, and 1:1, 3:4 and 4:3 for the rest. A live embed and a recorded clip can share a page if you size the frame to the same ratio.
Sources
Related posts
More in Sume Avatar 1.0
- Tavus Phoenix-4.5 adds cartoon and anime faces; Sume's options
Tavus Phoenix-4.5 supports cartoon, anime and Pixar-style faces from a photo or video. Sume creates avatars from a prompt, traits or a photo; test styles first.
- Two AI characters in one TikTok scene: one avatar per job, then cut
One Sume avatar job holds one avatar. For a two-person dialogue, make a job per speaker and alternate them in Timeline 1.0. Cost for 60 s of dialogue: $14.82.
- Viewers suspect AI within 20 seconds: disclose in the first frame
Tavus says Griffin-Lite doubters suspected within 20 seconds, and participants were told only afterward. Put the AI disclosure up front in a Sume clip.
- YouTube avatar vs Sume Avatar: selfie capture or prompt and photo
YouTube's avatar is made once from your own face and voice and used in its AI tools. How it is created, its limits, and how Sume's Avatar 1.0 differs.
Written by Sume