MiMo V2.6, Muse Spark 1.3 or Sonnet 5.5: Sume picker rows compared

Three catalog rows side by side, taken from Sume's own model catalog notes: MiMo V2.6 Pro, Muse Spark 1.3 and Sonnet 5.5. What is recorded, and what is not.

4 min readSume
All posts

For a video script, storyboard or caption job, Sume's catalog does not rank these three rows for those tasks; it ranks them by Artificial Analysis index, recorded in code comments as Sonnet 5.5 at 56.0, Muse Spark 1.3 at 48.1 and MiMo V2.6 Pro at 46.3. Those are general-intelligence scores, not video-writing scores. All three are image-and-attachment capable, thinking-enabled rows, so the practical way to choose is to run your own brief in each.

What do the rows have in common?

Each is a catalog entry with images, attachments and thinking enabled, routed through OpenRouter. MiMo and Muse Spark carry a "multimodal reasoning" tooltip; each of those two tooltips also states a 1M context. The MiMo tooltip adds a 128K maximum output. The Sonnet 5.5 row is Anthropic's official id claude-sonnet-5-5, per the comment in the source, dated 2026-09-28.

Three Sume picker rows, from catalog source notes read 2026-10-11
RowIndex score in code commentTooltip claimPicker note
Sonnet 5.556.0Anthropic row, id claude-sonnet-5-5Listed in the Anthropic group
Muse Spark 1.348.11M-context multimodal reasoningAfter the Anthropic and Google rows
MiMo V2.6 Pro46.31M context, 128K max output; multimodal reasoningAfter Grok 4.7

What is not recorded?

The catalog does not record a video-input flag for any of them, only images and attachments. It also records no benchmark for scripts, captions or shot lists. Vendors' own pages do not close that gap: Meta's Muse Spark 1.3 post describes agentic and coding gains and about 20% fewer tool calls, and Xiaomi's Hugging Face page lists parameter counts only.

How do I compare them for my job?

The API cannot do it, because Agent Completions accepts only sume-agent. Compare in the chat picker instead.

  • Write one brief, with the same assets attached, and send it to each row in a fresh chat.
  • Score the three outputs blind on length, caption line breaks and whether the shot list matches your clip durations.
  • Count tool calls before the first paid generation; stop each run with a small spending cap in the prompt.
  • Keep the winner for that format and re-test when the catalog changes.

Why not just pick the highest score?

The index scores in the code comments were recorded to order the picker, not to choose a writing model. Sonnet 5.5 tops the three on that index, but the index says nothing about how well a model fits a 9:16 caption layout or keeps to a 12-second clip budget.

There is also a budget angle. Each row has its own token rate in Sume's price book, and I did not find those rates stated on a page I could cite, so no cost comparison appears here. Run a small job in each row and read the usage view afterwards.

If your workspace's picker does not list all three, that is the OpenRouter catalog gate doing its job. You can only compare the rows you can see.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume