Ideogram 4 download: Hugging Face gate, login and first image
To run Ideogram 4 locally: accept the gate on Hugging Face, log in with hf, pip install the repo, run run_inference.py. The flags and the nf4 or fp8 choice.

Four steps, from Ideogram's model card: accept the gate on the Hugging Face page, authenticate with hf auth login or an HF_TOKEN variable, run pip install . in the inference repository the card links, and call python run_inference.py with a prompt, an output path and a quantization. Choose fp8 if you are not on an NVIDIA GPU, because the card says nf4 is CUDA only.
The weights are free for research and personal projects only; the licence is Ideogram 4 Non-Commercial. Everything below was read from Ideogram's Hugging Face pages on 2026-10-03, and I did not run the model myself.
Which build should you download?
The organisation page lists three Ideogram 4 repositories: ideogram-4-nf4, ideogram-4-fp8 and ideogram-4-nf4-diffusers. The fp8 card says it supports all hardware platforms and has no Diffusers support; the nf4 card says nf4 needs CUDA hardware. Neither card states a VRAM figure.
| Repository | Quantization | Hardware per the card | Diffusers |
|---|---|---|---|
| ideogram-4-fp8 | fp8 | All hardware | No |
| ideogram-4-nf4 | nf4 (4-bit) | CUDA only | Not stated on the card I read |
| ideogram-4-nf4-diffusers | nf4 (4-bit) | Not checked; listed on the organisation page | The name says Diffusers |
What do the steps look like?
Accepting the gate means agreeing on the model page to share your contact information. Then log in once and run the commands below from the inference repository. The flags are the ones the card shows: --prompt, --output, --quantization, --height, --width and --sampler-preset. For the highest quality it names V4_QUALITY_48 at 2048 by 2048.
- A Hugging Face account that has accepted the gate on the Ideogram 4 model page; without it the download is refused.
- A token: either the interactive
hf auth loginor an exportedHF_TOKEN. - One build chosen up front: fp8 for non-CUDA hardware, nf4 if you have a CUDA GPU and want the 4-bit file.
- An output path and a size that is a multiple of 16 between 256 and 2048 on each side.
- A decision on prompts: casual text through the magic-prompt step, or a JSON caption you write yourself.
hf auth login
pip install .
python run_inference.py \
--prompt "A bakery window with a hand-painted sign that reads OPEN AT SIX" \
--output out.png \
--quantization fp8 \
--height 1024 --width 1024
# highest quality, per the card:
# --height 2048 --width 2048 --sampler-preset V4_QUALITY_48Does it run fully offline?
Possibly not by default. The card's own example passes --magic-prompt-key "$IDEOGRAM_API_KEY", and it says the included command line uses a magic prompt step to turn a casual prompt into the structured JSON the model was trained on. I did not find a statement that this step runs locally. If you need an air-gapped run, write the JSON caption yourself and test with the key unset.
The model was trained exclusively on structured JSON captions, so plain text works but is not the format it prefers. The card's prompting guide covers the schema, colour palettes and bounding-box layout.
What about GPU memory?
Ideogram's cards give no number. A third-party deployment guide from Spheron reports roughly 8 to 10 GB of VRAM for nf4 and 12 to 15 GB for fp8; that is reported by Spheron, not stated by Ideogram, so confirm on your own card before you commit to hardware.
Resolution is a separate limit: the card says any size from 256 to 2048 pixels a side in multiples of 16, with aspect ratios up to 6:1.
How does Sume fit in?
Sume does not host Ideogram 4 and has no field for your own weights, so there is nothing to configure on the Sume side for a local run. Where Sume helps is the comparison and the next step. Its image catalog lists ideogram/ideogram-v3, which makes a useful hosted baseline for the same prompt, and the Video and Upscale routes take an image from your local run once you host the file at a public HTTPS URL; Sume's media input rules reject localhost, private-network and signed URLs.
Because the licence is non-commercial, keep anything you plan to publish for a client out of the local pipeline until you have a commercial tier. The licence post lists what each tier permits.
Sources
- Hugging Face: ideogram-ai/ideogram-4-fp8 model card (read 2026-10-03)
- Hugging Face: ideogram-ai/ideogram-4-nf4 model card (read 2026-10-03)
- Hugging Face: ideogram-ai organisation page (read 2026-10-03)
- Ideogram licensing page (read 2026-10-03)
- Spheron: Deploy Ideogram 4 on GPU Cloud (third-party, reported, read 2026-10-03)
- Image API
- Media inputs
Related posts
More in Developers
- Is AI avatar video real time? How long a Sume job takes
A Sume avatar video is a job, not a live stream: it queues, renders, and you poll or take a webhook. What the sync wait caps at, and a Python polling loop.
- Latin American Spanish text to speech: es or es-MX on Sume?
Sume's TTS language is a free string and its voice library tags voices with plain es. What that means for Mexican, Argentine or Spain Spanish, and how to test.
- Live AI avatar API: a Tavus conversation vs a Sume job
A live avatar API creates a room you join. Sume's Avatar API creates a job you poll. Field-by-field map of Tavus create conversation and Sume talking-video.
- MAI-Transcribe-2-Streaming Realtime API: events vs Sume job URLs
Microsoft's streaming transcriber uses a WebSocket with delta, intermediate and commit events. Sume STT takes a file URL and returns a job. A side-by-side.
Written by Sume