FLUX.2 [klein] needs about 13 GB VRAM: local run or API?

Recraft cites about 13 GB of VRAM for FLUX.2 [klein] under Apache 2.0. A checklist for deciding between a local GPU and a hosted image API.

4 min readSume
All posts

FLUX.2 [klein] needs about 13 GB of VRAM according to a Recraft post published October 3, 2026, and it ships under the Apache 2.0 license. Whether it runs well depends on your card and on what else shares it.

What is confirmed

The Black Forest Labs documentation lists open-weight FLUX.2 [klein] and [dev] on Hugging Face. The Recraft post adds the license and the VRAM figure for klein. Neither source gives throughput, so any images-per-minute number you see elsewhere should be measured on your own hardware.

FLUX.2 [klein] facts from the cited pages (read 2026-10-03)
ItemValueSource
LicenseApache 2.0Recraft post
VRAMAbout 13 GBRecraft post
WeightsOpen weights on Hugging Face (klein and dev)BFL documentation
ThroughputNot statedMeasure it yourself

Headroom math you should do before downloading

A 13 GB model on a 16 GB card leaves roughly 3 GB. That has to cover the operating system's display, the activations for your image size and any text encoder you keep resident. Raising the output resolution or the batch size grows activations, so the headroom you see at idle is not what you get under load.

A simple rule is to treat the quoted figure as a floor. Run one image at your target resolution, watch peak memory, then try your intended batch size. If you hit out-of-memory errors, you have your answer about which route to take.

When the API is the simpler route

A hosted API removes the GPU question entirely. On Sume, you send POST /v1/images and read the finished image from a Sume-hosted URL in data[].url. The call blocks for up to 30 seconds and then returns either 200 with the image or 202 with a job you can poll, so a slow render does not break your client. See Jobs and results for the polling loop.

Billing is per generated image, and failed generations are not billed. The endpoint pricing line in the catalog is the amount charged, so cost_usd × n is what a batch costs. Use GET /v1/images/models/{model_id}/endpoints to read the line for the model you pick.

  • Pick local when you own idle GPU time, need offline runs, or must keep inputs on your own machines.
  • Pick hosted when demand is bursty, when you want several model families behind one request shape, or when nobody on the team wants to maintain drivers.
  • Pick both when drafts can run locally and finals need a model you do not host.

Check before you commit

Run ten of your real prompts both ways, time them, and compare the results. Note the license for each model you use and keep a copy with your project files. The Apache 2.0 grant applies to klein; the other open models in this space carry their own terms.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume