Kimi K3 on Sume: no catalog row, and what to pick for a video agent
Kimi K3 is not in Sume's agent model catalog. Moonshot says it takes text, images and video; here is the verified alternative and how Kimi can still call Sume.

No. Kimi K3 is not an agent model on Sume: the catalog source has no Kimi or Moonshot row, and the Agent Completions API accepts only sume-agent. The nearest rows in the catalog are MiMo V2.6 Pro, Muse Spark 1.3 and Sonnet 5.5. If you want Kimi itself to drive Sume, run Kimi on your side and let it call Sume's hosted MCP tools.
What does Moonshot say about K3?
The Hugging Face model card lists 2.8 trillion total parameters, 104 billion active per token, a 1-million-token context window, and native text, image and video input, with a 401-million-parameter MoonViT-V2 vision encoder. The license is the Kimi K3 License, a custom license; the card tells commercial users to read its LICENSE file. The card says the API is reachable through platform.kimi.ai with OpenAI-compatible and Anthropic-compatible interfaces.
| Question | Kimi K3 (Moonshot model card) | Sume |
|---|---|---|
| Agent picker row | n/a | None in the catalog source |
| Context window | 1-million-token window | Not applicable |
| Input types | Text, images, video | Nearest rows take images and attachments |
| License | Kimi K3 License (custom) | Not applicable |
| Tools | Tool calls; pass assistant message back whole | MCP tools incl. jobs_wait, jobs_result |
Why does the tool-call note matter for media jobs?
The card says K3 "requires the complete assistant message returned by the API to be passed back to messages as-is", including reasoning_content and tool_calls. If you drive Sume's job tools from your own loop, that means appending the whole assistant message each turn, then appending the tool result from jobs_wait or jobs_result. Trimming the reasoning field to save tokens would break the contract Moonshot describes.
How can Kimi still use Sume?
Sume's MCP tool list is model-neutral: generate_video, generate_image, tts_create, stt_create, video_inspect, video_frames_create, plus the jobs_* family. Paid and write tools require an idempotency_key, and max_spend_usd and dry_run are available as safety controls. Connecting a Kimi client is covered in the related Kimi Code posts below.
The trade-off is plain: on Sume's own agent you get Sume's runtime, spend caps and job wiring. With Kimi in your own loop you carry that loop yourself.
What would I pick on Sume today for a video agent?
From the catalog source, the closest choices are MiMo V2.6 Pro and Flash, and Muse Spark 1.3 for a closed model with a 1M context. Anthropic's Sonnet 5.5 and Opus 5.5 rows are also there. None of them is Kimi, and none is a drop-in for Kimi's licence terms or its video input, so treat them as the nearest available rows, not equivalents.
If video input is the reason you wanted Kimi, note that Sume's catalog flags for the rows above say images and attachments only. For analysing a clip, extract stills with video_frames_create and send those, or read the transcript produced by stt_create.
If the reason was cost or open weights, run Kimi on your own infrastructure under its custom license and call Sume only for generation. That keeps Sume's role small: it creates the media and returns job results.
Sources
Related posts
More in Models
- Korean clip? Whistle's seven languages vs Sume STT
Cactus Whistle lists English, German, French, Spanish, Italian, Dutch and Polish. For a Korean clip use Sume STT with a ko hint, then caption it.
- Muse Spark 1.3 on Sume: picker row, tool use and media jobs
Meta says Muse Spark 1.3 uses about 20% fewer tool calls. Sume lists it as a catalog row behind the OpenRouter switch; the API cannot pick it.
- Nova 2.5 Sonic for a narration file? Speech-to-speech vs Sume TTS
Nova 2.5 Sonic is built for live voice agents. For a finished narration file, Sume TTS takes a transcript up to 20,000 characters and returns audio as a job.
- Qwen-Image-2.1-Turbo runs 8 steps; what does Sume's Qwen row expose?
Steps, CFG and seed are model-card settings. Sume's qwen/qwen-image row lists none of them: only ratio, n 1-4, references, output format. Check with one GET.
Written by Sume