Muse Spark 1.3 on Sume: picker row, tool use and media jobs

Meta says Muse Spark 1.3 uses about 20% fewer tool calls. Sume lists it as a catalog row behind the OpenRouter switch; the API cannot pick it.

4 min readSume
All posts

Sume's catalog has a Muse Spark 1.3 row, labelled "Meta Muse Spark 1.3, 1M-context multimodal reasoning". It is enabled, but it sits behind the OpenRouter catalog switch, and it is a chat-picker choice only: Agent Completions takes sume-agent and nothing else. Meta's own post says the model makes about 20% fewer tool calls than its predecessor, which matters if your agent loops over media jobs.

What did Meta actually announce?

Meta's research blog post, dated September 2, 2026, describes improved performance on agentic and coding tasks and says the model is available in Muse Code and the Meta Model API. It states roughly 20% fewer tool calls and 25% fewer tokens versus the previous version in coding work, and it describes the model using tools to build its own context. The page gave no price and no context window, so neither appears here as a Meta fact.

One caution on the tool-call claim: Meta measured coding tasks. It is not a measurement of video-job orchestration, so treat it as a reason to test, not a promise.

How does Sume list it?

Sume's catalog notes record it as a 1M-context multimodal row with an Artificial Analysis score of 48.1 in a code comment, which puts it above Grok 4.7 and MiMo V2.6 Pro in the picker order. Its capability flags are images, attachments and thinking.

Muse Spark 1.3: vendor statement vs Sume catalog, read 2026-10-11
ItemMeta's postSume catalog source
DateSeptember 2, 2026Added as a picker row
Where it runsMuse Code, Meta Model APIOpenRouter route, meta/muse-spark-1.3
Tool useAbout 20% fewer tool calls than prior versionTool calls go through Sume job tools
Context windowNot stated on the page1M, per the catalog tooltip
PriceNot disclosed on the pageNot repeated here

What would a media-job test look like?

Run the same brief in two picker rows and compare how many tool calls each makes before the first generate_video call. Use dry_run and a small max_spend_usd so the comparison costs nothing, then count calls in the transcript.

  • Same brief, same assets, one row each.
  • Count calls to video_inspect and video_frames_create before generation starts.
  • Compare only completed runs; a stalled run says nothing about call counts.

What should I not assume?

Do not assume that fewer tool calls means a cheaper media run. The expensive part of a Sume run is usually the generation job itself, which is billed by the media model, not by the chat model that asked for it. A chat model that makes fewer calls can still request one pricey clip too many, which is why the spending cap and the dry run exist.

Do not assume the row is visible. It is enabled in the catalog, but the OpenRouter switch decides whether the picker shows it. And do not assume the API route can reach it: Agent Completions documents one accepted model name, sume-agent, and states that other values return a 400.

Finally, Meta's page did not state a context window or a price, so any figure you see elsewhere is not from Meta's own post as I read it on 2026-10-11.

Sources

Related posts

More in Models

All Models posts

Written by Sume