Why did that stock avatar rank first? Catalog search score weights

Sume's avatar catalog search scores handle 8, keywords 7, search card 6, display name 5. Read search_reasons and fix a query that ranks the wrong presenter.

4 min readSume
All posts

When you search Sume's avatar catalog with a text query, each result carries a search_score and a search_reasons list that tell you why it ranked where it did. The score is a weighted sum over fields: handle 8, keywords 7, search card 6, display name 5, embedding text 3, best-for 2 and product categories 2, each multiplied by how well your query words matched that field.

This post reads those weights from the avatar search code on origin/main (2026-10-11) and from the OpenAPI schema for POST /v1/avatar-catalog/search. It explains how to read a surprising ranking and how to change the query or the filters to fix it. For how to browse without a query, see explore mode and seeds.

How is the lexical score computed?

Your query is lower-cased and split into tokens. Tokens shorter than 2 characters are dropped, and only the first 16 are used. For each scored field the code compares every token with the field text: an exact match of the whole field is worth 3, a token that matches a whole word in the field is worth 2, and a token found inside a longer word is worth 1.

That match value is multiplied by the field weight, and the products are added up. An avatar with profile metadata also starts at 1 point, one without metadata at 0. Fields that contributed at least once are returned in search_reasons, sorted alphabetically. Ties are broken by handle, then by id, so identical scores keep a stable order.

Lexical field weights for text search, from the Sume avatar search code (read 2026-10-11).
Field in search_reasonsWeightWhat it holds
handle8The avatar handle
keywords7Keyword list in the search metadata
search_card6Short text card for the avatar
display_name5Display name, or the name if none
embedding_text3Public-safe text prepared for embedding
best_for2Suitability list in the profile
product_categories2Suitability categories in the profile

A worked example with made-up fields

Take the one-word query skincare. Suppose avatar A lists skincare beauty routine as its keywords and mentions skincare as a whole word in its search card, so both score as whole-word matches (2 each): 2 x 7 + 2 x 6 = 26, plus 1 for having metadata, giving 27. A different avatar whose handle is exactly the query word scores 3 x 8 = 24 on that field alone, plus 1 for metadata, so 25: one exact handle hit lands just under two whole-word hits in keywords and the card.

The point of the arithmetic is the ranking logic, not the numbers. A single keyword hit is worth more than a search-card hit. If your query is a persona word like warm that appears in many search cards, those avatars tie on lexical score and the handle ordering decides, which can look arbitrary.

Why do the results sometimes reorder with an embedding reason?

Hybrid ranking is on by default. When a stored profile embedding and a query embedding are both available, the schema describes a weighted-sum fusion of the lexical score (0.65) and query cosine (0.35), and results that used it carry an embedding reason. If a vector is missing or the embedding provider fails, the request degrades to lexical plus facet diversity rather than failing.

Set hybrid: false to force the lexical path, which is the easiest way to debug a ranking: the order then follows the weights in the table. Use the default again once the query is fixed.

How do I steer a query that ranks the wrong presenter?

Put the discriminating word where the weight is. A product word is worth most when it is also in an avatar's keywords, so use the vocabulary of the catalog rather than marketing phrasing. Use filters for hard requirements: product_category, best_for and language only keep avatars whose profile metadata lists them, and an avatar with no metadata is dropped when one of those filters is set. status defaults to ready.

If too few results survive, auto_expand relaxes soft text first, then product_category, then best_for, and lists what it relaxed in relaxed_filters. Language, visibility and ownership are never relaxed. The order is walked through in the relaxed-filters post. Note that Avatar 1.0 talking videos are English-only, so a language filter shows catalog fit, not a promise of another spoken language.

curl -X POST https://api.sume.com/v1/avatar-catalog/search \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query": "skincare", "hybrid": false, "limit": 10}'

What to do with the chosen handle

Read the top result's handle, check metadata.suitability.avoid_for and brand_safety_notes as covered in casting notes for a stock avatar, then pass the handle as avatar_handle to POST /v1/avatar-1.0/talking-video, as described on Generate avatar video. A public catalog avatar needs no creation step, so the cost is the per-second video rate.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume