MiMo V2.6, Muse Spark 1.3 or Sonnet 5.5: Sume picker rows compared
Three catalog rows side by side, taken from Sume's own model catalog notes: MiMo V2.6 Pro, Muse Spark 1.3 and Sonnet 5.5. What is recorded, and what is not.

For a video script, storyboard or caption job, Sume's catalog does not rank these three rows for those tasks; it ranks them by Artificial Analysis index, recorded in code comments as Sonnet 5.5 at 56.0, Muse Spark 1.3 at 48.1 and MiMo V2.6 Pro at 46.3. Those are general-intelligence scores, not video-writing scores. All three are image-and-attachment capable, thinking-enabled rows, so the practical way to choose is to run your own brief in each.
What do the rows have in common?
Each is a catalog entry with images, attachments and thinking enabled, routed through OpenRouter. MiMo and Muse Spark carry a "multimodal reasoning" tooltip; each of those two tooltips also states a 1M context. The MiMo tooltip adds a 128K maximum output. The Sonnet 5.5 row is Anthropic's official id claude-sonnet-5-5, per the comment in the source, dated 2026-09-28.
| Row | Index score in code comment | Tooltip claim | Picker note |
|---|---|---|---|
| Sonnet 5.5 | 56.0 | Anthropic row, id claude-sonnet-5-5 | Listed in the Anthropic group |
| Muse Spark 1.3 | 48.1 | 1M-context multimodal reasoning | After the Anthropic and Google rows |
| MiMo V2.6 Pro | 46.3 | 1M context, 128K max output; multimodal reasoning | After Grok 4.7 |
What is not recorded?
The catalog does not record a video-input flag for any of them, only images and attachments. It also records no benchmark for scripts, captions or shot lists. Vendors' own pages do not close that gap: Meta's Muse Spark 1.3 post describes agentic and coding gains and about 20% fewer tool calls, and Xiaomi's Hugging Face page lists parameter counts only.
How do I compare them for my job?
The API cannot do it, because Agent Completions accepts only sume-agent. Compare in the chat picker instead.
- Write one brief, with the same assets attached, and send it to each row in a fresh chat.
- Score the three outputs blind on length, caption line breaks and whether the shot list matches your clip durations.
- Count tool calls before the first paid generation; stop each run with a small spending cap in the prompt.
- Keep the winner for that format and re-test when the catalog changes.
Why not just pick the highest score?
The index scores in the code comments were recorded to order the picker, not to choose a writing model. Sonnet 5.5 tops the three on that index, but the index says nothing about how well a model fits a 9:16 caption layout or keeps to a 12-second clip budget.
There is also a budget angle. Each row has its own token rate in Sume's price book, and I did not find those rates stated on a page I could cite, so no cost comparison appears here. Run a small job in each row and read the usage view afterwards.
If your workspace's picker does not list all three, that is the OpenRouter catalog gate doing its job. You can only compare the rows you can see.
Sources
Related posts
More in Comparisons
- Open video model licences, Oct 2026: Kandinsky, Prism, H3, LTX
Which of this month's open video models are MIT and which carry territory or revenue limits: Kandinsky 6.0, Prism, MiniMax H3 and FastH3 Trim, LTX-2.5.
- OpenRouter lists 26 video ids: which have a Sume id (Oct 2026)
Of 26 video ids on OpenRouter's listing, 11 map to a Sume catalog id and 15 do not. The full id-by-id table, read 2026-10-11, and what Sume lists instead.
- OpusClip API alternative: what Sume can cut and what it cannot
Sume has no highlight-finding API like OpusClip's. It trims, captions and assembles clips you pick; the source must be on media.sume.com first.
- PiAPI's $0.50 start credit and resetting bonus vs a Sume USD balance
PiAPI gives new accounts $0.50 plus bonus credits that reset each cycle. Sume bills a USD balance at published model rates and sets concurrency by plan.
Written by Sume