Anki cards with AI images: collection.media and the img tag
Anki finds images by filename in collection.media, via an img tag, with Allow HTML on. Make one PNG per word with Sume and write a tab-separated import file.

Put the image files in Anki's collection.media folder, write <img src="name.png"> in a field, and turn on Allow HTML when you import. The Anki manual's text-file import page says to copy the files into collection.media, not to use subdirectories, to make the field's filename match the file exactly, and to enable HTML in the import dialog so the tag renders.
The generating side is the Sume image API. One 200 response gives you data[0].url; download it and give it a name no other card will reuse.
A script for a word list
This Python script makes one small illustration per word, saves it with a sume- prefix, and writes cards.txt with a tab between the word and the image tag. Run it with a key in the environment, then copy the PNG files into collection.media.
import csv, json, os, urllib.request
KEY = os.environ.get("SUME_API_KEY", "")
if not KEY:
raise SystemExit("SUME_API_KEY is empty")
def post(body):
req = urllib.request.Request(
"https://api.sume.com/v1/images", json.dumps(body).encode(),
{"Authorization": "Bearer " + KEY, "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=60) as r:
return r.status, json.loads(r.read())
rows = []
for word in ["apple", "bridge"]:
prompt = f"simple flat illustration of a {word}, white background, no text"
status, out = post({"model": "openai/gpt-image-2.5", "prompt": prompt,
"quality": "low", "output_format": "png"})
if status != 200:
print(word, "status", status)
continue
name = f"sume-{word}.png"
urllib.request.urlretrieve(out["data"][0]["url"], name)
rows.append([word, f'<img src="{name}">'])
with open("cards.txt", "w", newline="") as f:
csv.writer(f, delimiter="\t").writerows(rows)What does the manual warn about?
Anki's page says embedding image references directly in card templates is not supported, and tells you to put them in fields instead; the script follows that by writing the tag into the second column. Because the filename must match exactly, keep the sume- prefix and avoid spaces in words that you turn into file names.
Write no text in the prompt. Image models can misspell lettering, and a wrong spelling on a vocabulary card teaches the wrong thing. Review each picture before you study from it.
A 202 means the job outlived the 30-second wait. The script prints the status and skips that word, so rerun later or poll the job as the jobs guide describes. Each success bills one image, and usage.cost has the USD amount if you want to total a deck.
For a deck of a few hundred words, run the loop in batches and keep a log of word, file name and cost. The request accepts n from 1 to 10 where the model allows it, but a flash card needs one picture per word, so n of 1 is the right setting here. If a picture misses, change the prompt and regenerate under the same file name, then replace the file in collection.media; the field text does not change, because the name stays the same.
Sources
Related posts
More in Use cases
- Annual compliance refresher as an avatar video: review and records
An avatar can deliver a 45-second compliance refresher. Legal approves the script; you keep the video id and transcript as a record of what was said.
- App demo promo clip: generate the scene, composite the real screen
Don't ask a video model to invent your app UI. Generate the scene with Sume, then stack your real screen recording with Timeline compose, $0.02 flat per shot.
- App screenshots to a 30-second product demo with Wan 3.0 on Sume
Turn six app screenshots into a 30 second demo on Sume: five 6 second first-and-last-frame Wan 3.0 clips cost $3.75 at 720p. Setup, limits, and what to check.
- Product demo video from a screen recording: $0.34 on Sume
Cut a 3-minute screen recording to a 30-second demo, add a voiceover and captions on Sume for about $0.34. Google Play autoplays only the first 30 seconds.
Written by Sume