Ming-Image open weights or a hosted image API: what you operate
Ming-Image-0.1-Design is MIT-licensed and runs on one 80 GiB GPU. This post lists what self-hosting means in practice and what a hosted API call replaces.

Open weights mean you run the model; a hosted API means you send a request. inclusionAI's Ming-Image-0.1-Design and Ming-Image-0.1-Design-Layer, released 2026-09-17, are two 6B-parameter MIT-licensed models; the Layer model outputs RGBA layers and the pair is reported to run on a single 80 GiB GPU (digitalapplied tracker, read 2026-10-02).
Sume is the hosted side. It serves the image models in its catalog through one API and does not host Ming weights or give you weights to run, per the docs I read on 2026-10-02.
What do the Ming models do?
Per the tracker, Design is a design-focused image model and Design-Layer decomposes a flat design into editable layers with transparent RGBA outputs. Both are MIT licensed. I read only the tracker, not the model cards, so check the licence and hardware notes on the model page before you plan around them.
| Item | Reported fact |
|---|---|
| Models | Ming-Image-0.1-Design and Ming-Image-0.1-Design-Layer |
| Size | 6B parameters each |
| Licence | MIT |
| Output | RGBA transparent outputs; layer decomposition |
| Hardware | One 80 GiB GPU |
| Date | 2026-09-17 |
What does a hosted call replace?
You skip the GPU, the model server, and the update work. What you take on instead is the request contract: Sume jobs are durable, a 202 job envelope arrives when a generation outlasts the 30-second sync wait, and image generation bills only when it completes.
Results come back as Sume-hosted signed URLs under data[].url, so you store the Sume URL rather than a raw provider URL.
Which one fits which job?
Choose by the output you need. If you need true editable layers today, a layer model is the match, and Sume's route is a workaround: separate calls per layer, as in editable layers with Pillow. If you need flat images, edits, references and masks behind one key and one billing line, the Image API is that. Pilot both on ten real designs before you commit.
Sources
Related posts
More in Comparisons
- MiniMax Speech 2.8 pitch and emotions vs Sume TTS controls
MiniMax T2A lists speech-2.8 models, nine emotions, pitch -12 to 12 and 10,000 characters. Sume TTS has speed, volume and emotion text, no pitch.
- MiniMax <#1.5#> pause markup vs Sume TTS: no pause marker
MiniMax inserts pauses with <#x#> markers from 0.01 to 99.99 seconds. Sume TTS documents no pause marker; here is how to split a script and what to use instead.
- MiniMax Video Agent template API: submit, poll, download
MiniMax Video Agent builds a video from a template_id plus your media and text. The endpoints, statuses, and what Sume offers for template-driven video.
- Mirage Tesseract free local engine vs a hosted avatar video API
Mirage Tesseract is an agent video suite with a free local engine. When does a hosted avatar API like Sume Avatar 1.0 fit better? Differences, with dated facts.
Written by Sume