Anam Cara-4 Director Notes vs Sume scene prompts and silence
Anam Director Notes steer a live avatar with a style and an Expressivity control. Sume directs scenes with a prompt or photo, silence beats and a quality tier.

Sume has no counterpart to Director Notes or an Expressivity dial in the docs cited here. What a Sume request can direct is the scene (a prompt or a photo), the pacing (silence beats and duration) and the quality tier.
Anam facts are from its Cara-4 post; Sume facts from Generate avatar video, read 2026-10-01.
What are Director Notes?
Anam says Director Notes, in beta with Cara-4, set the performance with a preset style (such as warm, supportive, angry or distressed) or a short custom prompt, and an Expressivity control sets how strongly the model follows it. During a session the LLM can add performance cues as it streams the spoken response. Cara-4 is available through the API with avatarModel: "cara-4".
What can I direct in a Sume avatar video?
scene takes { "type": "prompt", "prompt": "..." } or { "type": "photo", "image_url": "https://..." } for scene direction. Multi-scene video_inputs suit hooks, demos or silence beats. A voice.type: "silence" scene is a non-speaking beat and needs a duration. quality is standard, plus (default) or max; max is the highest tier with slower turnaround.
| Control | Anam Cara-4 | Sume avatar video |
|---|---|---|
| Performance style | Preset or prompt, plus Expressivity | Not documented |
| Scene | Not in the post | scene prompt or photo |
| Pacing | LLM cues mid-session | Silence beats, duration |
| Tier | Model choice | standard, plus, max |
Does this change when the output is live or rendered?
Yes. Anam's controls act during a live session. Sume renders a finished clip from a script you wrote, so direction is settled before submit. For voice pace and mood, see avatar emotion and speaking speed.
Sources
Related posts
More in Models
- Anam Cara-4 portrait 768x1152 vs Sume 9:16 avatar at 720p
Anam Cara-4 renders natively at 768x1152 portrait. Sume's avatar video defaults to 9:16 at a fixed 720p, with plans estimated at 4-60 seconds.
- Dictation API: AssemblyAI cleaned text vs Sume STT word timings
AssemblyAI's Dictation API returns send-ready text. Sume STT returns a transcript with words[] timings and no cleanup flag, so you do the filler removal.
- AssemblyAI text to speech: coming soon, and what Sume has today
AssemblyAI's product menu lists a Text-to-Speech API as coming soon. Sume TTS is available now: what a call takes, the 20000-character cap, and word timings.
- AssemblyAI Universal-3.6 Pro 32 languages vs Sume STT language_code
AssemblyAI streaming defaults to Universal-3.6 Pro with 32 languages; 3.5 Pro has 19. Sume STT 1.0 takes one optional language_code hint and auto-detects.
Written by Sume