MiMo V2.6 Pro and Flash in Sume's agent picker: what the catalog says

Sume's model catalog has rows for Xiaomi MiMo V2.6 Pro and Flash behind the OpenRouter switch. What the repo records, and what the API cannot pick.

4 min readSume
All posts

Yes, with a caveat: Sume's model catalog contains two Xiaomi rows, MiMo V2.6 Pro and MiMo V2.6 Flash, and both are marked enabled. They are gated behind Sume's OpenRouter catalog switch, so whether you see them depends on the picker in your workspace. They are chat-picker rows only. The Agent Completions API accepts one model name, sume-agent, and answers any other value with a 400.

What does the Sume catalog record for the two MiMo rows?

The rows were added as catalog entries for Xiaomi's V2.6 pair. Each carries a tooltip that reads "1M context, 128K max output; multimodal reasoning". The capability flags on both are images, attachments and thinking turned on. The routing target is the OpenRouter ids xiaomi/mimo-v2.6-pro and xiaomi/mimo-v2.6-flash.

The catalog notes also record a picker position for each: Pro sits after Grok 4.7 with an Artificial Analysis score of 46.3 noted in a code comment, and Flash follows it with no score because it was not on the index when the row was added.

MiMo V2.6 rows in Sume's catalog source, read 2026-10-11
RowTooltip in catalogCapability flagsGate
MiMo V2.6 Pro1M context, 128K max output; multimodal reasoningimages, attachments, thinkingopenrouter
MiMo V2.6 Flash1M context, 128K max output; multimodal reasoningimages, attachments, thinkingopenrouter

What does Xiaomi itself list?

Xiaomi's Hugging Face organization page lists MiMo-V2.6-Pro at 1T parameters and MiMo-V2.6-Flash at 311B, each with reinforcement-learning variants. That page, as I read it, gave no context length or license text, so this post does not repeat either number as a Xiaomi claim. The 1M context figure above is Sume's own catalog wording.

Can a MiMo row run a media job?

A picker row only changes which language model drives the chat. The media work still goes through Sume's job tools, the same ones any agent sees: generate_video, generate_image, tts_create, stt_create, video_frames_create, with jobs_wait and jobs_result to collect output. Paid tools need an idempotency_key, and max_spend_usd and dry_run are optional controls you can add.

So a MiMo row can plan shots, write captions and call those tools. It cannot generate video itself, and the catalog flags above say nothing about video input: only images and attachments.

How do I check my own picker?

Do not assume an enabled row means a visible row. The gate decides that.

  • Open the agent model picker and look for MiMo V2.6 Pro or Flash. If neither is listed, the OpenRouter catalog is off for you.
  • If you are calling the API, do not send a model name other than sume-agent.
  • If you need a row today, the catalog tests list Opus 5.5 on the closed picker alongside Auto, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Sol.

Pro or Flash for a script-and-caption job?

The catalog does not rank them for writing tasks. It only records that Pro has an Artificial Analysis score of 46.3 in a code comment and Flash has none. That is a general-intelligence index, so it says little about caption timing or shot lists. A fair test is to send the same brief to both rows and compare the output on your own clip lengths.

Expect the sizes Xiaomi lists on Hugging Face, 1T against 311B, to matter for quality and speed, but do not assume a direction without running your brief. Neither the catalog nor the Hugging Face page gives a price per token, so cost comparisons here would be invention; check the usage view in your workspace after a run instead.

If you only need a quick rewrite of a hook line or a set of titles, the smaller row is the natural first try. If the job has many steps, such as inspecting a clip, cutting frames, then planning regeneration, run the larger row on a copy of the job first.

Sources

Related posts

More in Models

All Models posts

Written by Sume