FLUX.2 [klein] needs about 13 GB VRAM: local run or API?
Recraft cites about 13 GB of VRAM for FLUX.2 [klein] under Apache 2.0. A checklist for deciding between a local GPU and a hosted image API.

FLUX.2 [klein] needs about 13 GB of VRAM according to a Recraft post published October 3, 2026, and it ships under the Apache 2.0 license. Whether it runs well depends on your card and on what else shares it.
What is confirmed
The Black Forest Labs documentation lists open-weight FLUX.2 [klein] and [dev] on Hugging Face. The Recraft post adds the license and the VRAM figure for klein. Neither source gives throughput, so any images-per-minute number you see elsewhere should be measured on your own hardware.
| Item | Value | Source |
|---|---|---|
| License | Apache 2.0 | Recraft post |
| VRAM | About 13 GB | Recraft post |
| Weights | Open weights on Hugging Face (klein and dev) | BFL documentation |
| Throughput | Not stated | Measure it yourself |
Headroom math you should do before downloading
A 13 GB model on a 16 GB card leaves roughly 3 GB. That has to cover the operating system's display, the activations for your image size and any text encoder you keep resident. Raising the output resolution or the batch size grows activations, so the headroom you see at idle is not what you get under load.
A simple rule is to treat the quoted figure as a floor. Run one image at your target resolution, watch peak memory, then try your intended batch size. If you hit out-of-memory errors, you have your answer about which route to take.
When the API is the simpler route
A hosted API removes the GPU question entirely. On Sume, you send POST /v1/images and read the finished image from a Sume-hosted URL in data[].url. The call blocks for up to 30 seconds and then returns either 200 with the image or 202 with a job you can poll, so a slow render does not break your client. See Jobs and results for the polling loop.
Billing is per generated image, and failed generations are not billed. The endpoint pricing line in the catalog is the amount charged, so cost_usd × n is what a batch costs. Use GET /v1/images/models/{model_id}/endpoints to read the line for the model you pick.
- Pick local when you own idle GPU time, need offline runs, or must keep inputs on your own machines.
- Pick hosted when demand is bursty, when you want several model families behind one request shape, or when nobody on the team wants to maintain drivers.
- Pick both when drafts can run locally and finals need a model you do not host.
Check before you commit
Run ten of your real prompts both ways, time them, and compare the results. Note the license for each model you use and keep a copy with your project files. The Apache 2.0 grant applies to klein; the other open models in this space carry their own terms.
Sources
Related posts
More in Comparisons
- Gemini API webhooks for batch jobs: keep a poll backup, as on Sume
Gemini API release notes say webhooks replace polling for Batch and long-running operations. What changes in a client, and why Sume keeps a polling backup.
- Gemini Omni Flash 1,240 vs Seedance 2.0 1,225: what 15 Elo means
Hedra lists Gemini Omni Flash first at 1,240 Elo and Seedance 2.0 4K second at 1,225. A 15-point gap is about a 52% win rate; check limits before choosing.
- Google's video default is Omni Flash: three cases where Veo 3.1 fits
Google's video docs call Gemini Omni Flash the default and keep Veo 3.1 for extension, frame-specific and image-directed jobs. What each maps to on Sume.
- Google Ads built-in image and video generation vs a Sume pipeline
Demand Gen can generate up to 20 images per prompt and auto-build video from a logo, two images and two texts. Where a separate Sume pipeline differs.
Written by Sume