Let an agent pick talking video, Fabric or H3 Max over MCP

Rules for an agent on hosted MCP: avatar-videos_create for scripts, avatar-image-to-video_create for audio, and dry_run plus max_spend_usd before any paid call.

4 min readSume
All posts

Give an agent one rule per tool. Use avatar-videos_create when you have a script and want Sume to make the speech and the video. Use avatar-image-to-video_create when you already have an audio file and a still. Choose between VEED Fabric and MiniMax H3 Max Lip Sync with the model field on that second tool, and always run dry_run first. Sume avatars are async jobs: you submit a request, get a job id back, and read the finished video later. There is no live video session.

These are tools on Sume's hosted MCP server. They create jobs, so an agent still has to wait with jobs_wait and read jobs_result.

Which tool for which input

Avatar tools on hosted MCP (Sume docs, read 2026-10-07)
You haveToolNotes
A script and a ready avataravatar-videos_create4 to 60 seconds; quality standard, plus or max
A script and no avatar yetavatars_create, then avatar-videos_createAvatar creation is a separate paid job
A still and an audio fileavatar-image-to-video_createmodel veed/fabric-1.0 by default
A still and 5 to 14.8 s of audioavatar-image-to-video_create with model minimax/h3-max/lip-syncThe explicit alternative
First-frame stills to approveavatar-video-previews_createThen _generate_video
A lip-sync id on generate_videoNoneRefused with wrong_tool

Dry run and a spend cap

The paid-tool playbook says to call with dry_run: true first and examine the estimate, the balance and the queue behavior, then submit again with dry_run omitted or false. Add max_spend_usd to set a cap, and always pass an idempotency_key. An agent that follows this order cannot spend money on a surprise, because it sees the price before it commits.

A good instruction is short: "Before any paid avatar tool, dry-run it, show me the estimate, and wait for my yes." For an unattended agent, replace the last clause with a cap you set in advance.

{
  "idempotency_key": "avatar-video-2026-10-07-001",
  "dry_run": true,
  "max_spend_usd": 5,
  "payload": {
    "avatar_handle": "launch_host",
    "script": "Here is what is new this week.",
    "quality": "standard"
  }
}

Rules for the prompt

  • Use the avatar handle only when the user named that avatar; otherwise use a generated, inspected still.
  • Talking shots use text to speech first and then a lip-sync model, never a plain video model with narration.
  • Use one lip-sync model for each video, and measure the first clip's frame rate.
  • Audio outside 5 to 14.8 seconds goes to Fabric, never to a re-cut file.
  • After submit, wait with jobs_wait in slices and retry the wait on wait_slice_expired; never submit the paid call again.

What the agent should report

Ask the agent to return the job id, the final status, the media URL from the result, and the cost estimate it saw in the dry run. With those four items you can check the work without re-reading the chat.

Waiting correctly

The hosted MCP jobs_wait call holds for at most 55 seconds on each call, with a default of 50. A render takes longer than that, so the agent must expect a wait_slice_expired and call jobs_wait again with the same job id. It can also wait on up to 20 ids in one call with job_ids and wait_for set to all or any. With any, the other jobs still run and still bill.

A 524, 522, 523 or 525 on a wait is a transport failure and never a job outcome. The agent should wait again on the same ids, or read jobs_status once. It must not submit the paid create again, and it must not report the job as blocked.

A worked order of operations

Say a user asks for a 12-second product clip with a spokesperson. The agent lists avatars, finds one the user named, and calls avatar-videos_create with dry_run: true, a script, quality: "plus", and a max_spend_usd of 5. The estimate comes back near $2.94 at the no-product rate, or $3.096 with a product image. The agent shows the number, the user agrees, and the agent repeats the call without dry_run. Then it waits in slices, reads the result, and reports the media URL. If the user instead brings a recorded 9-second voice line, the agent picks avatar-image-to-video_create and checks the 5 to 14.8 second window before it chooses between the two lip-sync models.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume