MCP avatar-image-to-video_create: choose Fabric or H3 Max
On Sume's hosted MCP, one tool, avatar-image-to-video_create, makes still-plus-audio talking clips. A model field picks Fabric or MiniMax H3 Max lip sync.
On Sume's hosted MCP server, avatar-image-to-video_create is the one paid tool for a talking clip made from a still and an audio file. Its model field selects veed/fabric-1.0, the default, or minimax/h3-max/lip-sync. Use generate_video for anything else; it refuses the lip-sync id with wrong_tool.
This post covers when an agent should choose each model and what the call looks like. It uses Sume's MCP tools and gates and models docs.
What does the tool take?
The same body as the HTTP routes, inside payload: audio_url, a measured duration_seconds, and exactly one visual source, either an image_url or an avatar_handle. The tool is listed under Avatars and talking-head in the hosted inventory and is paid, so like every paid tool it needs an idempotency_key. dry_run=true previews cost without submitting, and max_spend_usd caps a call when you provide it.
Sume's guidance for agents is to prefer the generated, inspected posed still as image_url, and to use avatar_handle only when the user named that avatar.
{
"idempotency_key": "talking-clip-2026-10-03-001",
"dry_run": true,
"payload": {
"model": "minimax/h3-max/lip-sync",
"image_url": "https://example.com/host-still.png",
"audio_url": "https://media.sume.com/example/line.mp3",
"duration_seconds": 8.2,
"resolution": "768p"
}
}When should the agent pick H3 Max?
Only when the audio fits its window, 5 to 14.8 seconds. Fabric stays the default for everything else, including lines shorter than 5 seconds and longer ones, because the image-to-video body accepts a duration_seconds of 1 to 300.
Sume refuses a duration outside the H3 Max window rather than clamping it, since a clamped reservation would render a clip shorter than the audio and drop speech. Never re-cut audio to fit a model; send the clip that needs Fabric to Fabric.
| Question | Fabric (default) | H3 Max lip sync |
|---|---|---|
model value | veed/fabric-1.0 | minimax/h3-max/lip-sync |
| Audio length | 1-300 s duration_seconds on the image-to-video body | 5-14.8 s |
| Visual source | image_url or avatar_handle, not both | Same |
| Wrong tool | Not for generate_video | generate_video returns wrong_tool |
| Price basis | Per-second Fabric rate | Provider list per second times 1.25 |
What should the agent do before and after the call?
Before: tools_schema for the tool name, then a dry_run and a look at balance and queue. After: poll with jobs_status or jobs_wait and read jobs_result. jobs_wait holds a single call for a limited time, so treat an expired slice as a reason to wait again, not to submit again.
In reports, use Sume public ids and media.sume.com URLs, and never paste signed URLs, OAuth tokens or API keys into chat logs. If a session has only OAuth mcp:read, mutating and paid tools are hidden and calls return insufficient_scope.
How do I write the agent instruction?
Give the agent the rule, not the model name. For example: measure the audio first; if it is 5 to 14.8 seconds, payload.model may be minimax/h3-max/lip-sync; otherwise omit model and Fabric is used. Always call tools_schema with name: "avatar-image-to-video_create" once per session to read the live contract, because the hosted catalog, not this post, is the source of truth.
Keep the Agents @ catalog in mind too: the lip sync model appears there as @minimax-h3-max-lip-sync, next to MiniMax H3 Max, so a user can name it explicitly. When they do, honor it, then check the window before submitting.
Sources
Related posts
More in Sume Avatar 1.0
- Recurring AI lead for a TikTok series: $0.95 once, then per episode
Make a series lead once with Sume Avatar 1.0 for $0.95, then bill each episode by the second. 60, 180 and 300 second costs at three quality tiers.
- Shortest AI avatar video: 4 seconds, from $0.74 on Sume Standard
Sume avatar videos run 4 to 60 seconds. Per-second rates for Standard, Plus and Max, the 4 second floor, and what 15, 30 and 60 second clips cost.
- Synthesia's hosted avatar pipeline (Oct 1) vs Sume clips
Synthesia's no-code hosted LLM, TTS and avatar pipeline was due Oct 1, 2026 for live conversation. If you need a rendered clip instead, Sume takes a script.
- Tavus Phoenix-4.5 adds cartoon and anime faces; Sume's options
Tavus Phoenix-4.5 supports cartoon, anime and Pixar-style faces from a photo or video. Sume creates avatars from a prompt, traits or a photo; test styles first.
Written by Sume