Tavus Memory Stores vs Sume: personalizing avatar video per person

Tavus PALs now keep persistent memory per participant. Sume avatar videos are one-shot renders, so personalization is in the script you send. Here is the split.

5 min readSume
All posts

Tavus added Memory Stores on September 16, 2026: each PAL gets persistent memory per participant, with automatic profile learning and optional pinned memories. Sume has no equivalent, because a Sume avatar video is a rendered, script-driven clip, not a live conversation. If you want a video that feels personal to each viewer, you put what you know about that person into the script before you submit.

The Tavus detail is from its changelog, read on 2026-10-02. Sume's side is from Generate avatar video and Jobs and results.

What does a memory store do that a rendered video cannot?

Per the changelog, the memory lives with the PAL and is kept per participant, so a returning person is recognized across conversations. That is a property of a live agent: it listens, remembers and answers. A Sume avatar job takes a request in and returns a finished MP4; nothing in it listens or persists between runs. For the live-versus-rendered split, see real-time AI avatar vs video avatar API.

How do you personalize a Sume avatar video per person?

Keep the memory in your own system and render from it. A CRM row becomes a script, the script becomes one submit, and an Idempotency-Key derived from that row lets you retry the submit with the same key.

The same avatar handle is reused for every recipient; only the script changes. Mind the window: the estimated duration must be 4 to 60 seconds, so a short opener plus one specific line fits and a long recap does not.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: followup-crm-row-48213" \
  -d '{
    "avatar_handle": "studio_presenter",
    "aspect_ratio": "9:16",
    "quality": "standard",
    "script": "Hi Dana, thanks for the call about the Q4 rollout. Here is the next step we agreed.",
    "mode": "async"
  }'

What stays manual on the Sume side?

Four things are yours to handle.

  • Remembering who watched what: Sume does not store viewer history; keep it in your database.
  • Choosing the opener: there is no generated greeting; write the first line, or the first scene of video_inputs, yourself.
  • Approving frames: use an avatar video preview to review first-frame stills before the full render.
  • Delivery: poll the job or send mode: "webhook" with a webhook_url and read the result URL.

Which should you choose?

If the viewer needs to ask follow-up questions and be remembered, that is a live agent, and Tavus's page documents that product. If the viewer needs a short personal clip in an email or a message, render it: one job per recipient, the same avatar, a different script. For the broader difference between the two, read AI avatar vs AI agent.

What are the privacy consequences of each design?

A memory store means the vendor keeps profile information about the people who talk to the agent, so you owe those people a clear explanation and a way to remove it. With the Sume pattern the personal data lives in your CRM, and only the sentence you choose to put in the script reaches the render request. That is a smaller footprint, but it is still a script about a real person: do not include anything in it you would not say to them in an email.

Also remember the viewer is watching a synthetic presenter. Disclosure rules differ by platform and country, and a one-to-one video does not exempt you.

What does a minimal pipeline look like?

The pipeline has four steps. Select the rows that need a follow-up. Build a script per row, kept short enough to land inside the 4 to 60 second window. Submit one job per row with an idempotency key derived from the row id. Receive the terminal event on a webhook or poll the job, and store the returned media.sume.com URL against the row.

Queue capacity applies when you submit many at once; submit responses include generation_limits and a full queue returns 429 queue_full. For batches, read submitting 20 avatar videos at once.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume