Tavus Griffin single photo input vs a Sume photo avatar handle
Griffin works from one reference photo. In Sume you turn one photo into a named avatar handle once, then reuse it in every script. Steps, inputs and cost.
Yes, Sume can start from a single photo too, but it works differently from Griffin. Tavus says Griffin generates a live, conversational video from one reference photograph, and the model is a research preview for select testers. In Sume, one photo becomes a reusable avatar through POST /v1/avatar-1.0/generate for $0.95 flat, and every later clip references that avatar by handle.
Create the avatar from a photo
The creation request takes avatar_handle and an input union. The three forms are a prompt, props (ethnicity, sex and age) or a photo object with an image_url. A leading @ in the handle is stored without it.
Media inputs must be public HTTPS URLs. The media input rules reject localhost, private-network, signed or private URLs and mismatched content types before submission, so host the photo somewhere public and with the right content type.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"avatar_handle": "product_host",
"input": {
"photo": { "image_url": "https://example.com/host.jpg" }
}
}'
What each side asks of you
Tavus's page frames the single photo as the entire input to a live model. The Sume flow adds an explicit step, which costs once and pays back across clips: one handle, many scripts, one consistent face.
| Fact | Value |
|---|---|
| Tavus Griffin input | A single reference photograph (research preview) |
| Sume avatar creation | Prompt, props or photo, $0.95 flat |
| Reuse | Avatar handle across avatar videos |
| Customer availability, Griffin | Not available at this time |
Then render with the handle
Once the avatar is ready, an avatar video request names the handle and gives exactly one of script or video_inputs. Pricing per second depends on quality and whether a product_image is supplied. At standard, the rate without a product is $0.184 per second, so a 10 second clip is $1.84 on the Sume API pricing page.
Preview the first frame before you commit to a full render if the photo is a close call.
Sources
Related posts
More in Sume Avatar 1.0
- Tavus Memory Stores vs Sume: personalizing avatar video per person
Tavus PALs now keep persistent memory per participant. Sume avatar videos are one-shot renders, so personalization is in the script you send. Here is the split.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume