Anki cards with AI images: collection.media and the img tag

Anki finds images by filename in collection.media, via an img tag, with Allow HTML on. Make one PNG per word with Sume and write a tab-separated import file.

5 min readSume
All posts

Put the image files in Anki's collection.media folder, write <img src="name.png"> in a field, and turn on Allow HTML when you import. The Anki manual's text-file import page says to copy the files into collection.media, not to use subdirectories, to make the field's filename match the file exactly, and to enable HTML in the import dialog so the tag renders.

The generating side is the Sume image API. One 200 response gives you data[0].url; download it and give it a name no other card will reuse.

A script for a word list

This Python script makes one small illustration per word, saves it with a sume- prefix, and writes cards.txt with a tab between the word and the image tag. Run it with a key in the environment, then copy the PNG files into collection.media.

import csv, json, os, urllib.request

KEY = os.environ.get("SUME_API_KEY", "")
if not KEY:
    raise SystemExit("SUME_API_KEY is empty")

def post(body):
    req = urllib.request.Request(
        "https://api.sume.com/v1/images", json.dumps(body).encode(),
        {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=60) as r:
        return r.status, json.loads(r.read())

rows = []
for word in ["apple", "bridge"]:
    prompt = f"simple flat illustration of a {word}, white background, no text"
    status, out = post({"model": "openai/gpt-image-2.5", "prompt": prompt,
                        "quality": "low", "output_format": "png"})
    if status != 200:
        print(word, "status", status)
        continue
    name = f"sume-{word}.png"
    urllib.request.urlretrieve(out["data"][0]["url"], name)
    rows.append([word, f'<img src="{name}">'])
with open("cards.txt", "w", newline="") as f:
    csv.writer(f, delimiter="\t").writerows(rows)

What does the manual warn about?

Anki's page says embedding image references directly in card templates is not supported, and tells you to put them in fields instead; the script follows that by writing the tag into the second column. Because the filename must match exactly, keep the sume- prefix and avoid spaces in words that you turn into file names.

Write no text in the prompt. Image models can misspell lettering, and a wrong spelling on a vocabulary card teaches the wrong thing. Review each picture before you study from it.

A 202 means the job outlived the 30-second wait. The script prints the status and skips that word, so rerun later or poll the job as the jobs guide describes. Each success bills one image, and usage.cost has the USD amount if you want to total a deck.

For a deck of a few hundred words, run the loop in batches and keep a log of word, file name and cost. The request accepts n from 1 to 10 where the model allows it, but a flash card needs one picture per word, so n of 1 is the right setting here. If a picture misses, change the prompt and regenerate under the same file name, then replace the file in collection.media; the field text does not change, because the name stays the same.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume