avatars_search: hybrid by default, hybrid=false for exact handles

Sume avatars_search defaults to hybrid ranking for semantic queries. Send hybrid=false to match an exact handle or keyword, and filter for ready avatars.

5 min readSume
All posts

The Sume avatars_search tool ranks with a hybrid method by default, which suits a described person such as a friendly presenter in her thirties. When you know a handle or exact keyword, send hybrid=false to switch to lexical and facet matching only. It is a read tool, so it is available on an OAuth mcp:read session and costs nothing to try.

This is from the avatars_search tool contract in the Sume MCP server and from MCP tools and gates, read on 2026-10-03.

What the tool searches and returns

The tool calls the avatar catalog search endpoint and returns reusable avatar identity images. Each result exposes a preview_image_url and public-safe metadata. The handles identify saved looks; they are not video production inputs, and the catalog still must never be used as the talking-head plate.

Discovery is optional. An ordinary presenter or voice request does not need catalog casting at all, and an avatar the user already @-mentioned arrives with its own image URL, so there is no search and no recast.

avatars_search settings (read 2026-10-03)
SettingEffectUse when
hybrid omitted or trueHybrid ranking, the API defaultSemantic, descriptive queries
hybrid=falseLexical and facet matching only (no embedding fusion)You know the handle or exact term
filters.status=readyOnly ready avatarsAlmost always
Language filtersEnglish or Korean display namesFiltering by language

Read the images, not just the ranking

The tool guidance is to read the returned image or images and any available look notes rather than trusting a ranking score. A hybrid score tells you how close the text matched, not whether the face suits the shot. Have the agent open the top few preview images before choosing.

What happens after a pick

The catalog pick is only an identity reference. The route for a requested video is to inspect the chosen image, then generate a still with that image as an input reference for this shot's pose and framing, inspect and accept the new still, and then for a speaking shot create the voice track and run the image-to-video avatar tool on that still. A wordless beat uses the video generation tool with the accepted stills instead.

That keeps spend in the later, explicit paid steps. The search itself is a read, so an agent can run several searches, including one with hybrid=false, before it commits to anything paid.

Casting without overspending

A cheap casting loop has three steps. Search with a descriptive query and filters.status=ready. Open the top few preview images and note which look suits the shot. If the user named a specific handle, repeat with hybrid=false to confirm it exists exactly as written.

None of this spends money. The paid steps come only after the agent has a chosen identity image and the user has asked for a video. If a dry run or preview is available for the paid tool, run it before the first real submit, and pass max_spend_usd so the request carries its own ceiling.

When not to search at all

Skip the catalog when the user supplied a face, when the request is a plain voiceover, or when an @-mentioned avatar is already in the turn. Searching anyway adds a step and can swap the person the user asked for. The tool guidance is explicit that catalog casting is optional discovery, not a required stage.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume