Ideogram 4.0 claims 0.97 OCR accuracy: measure your own text score
Ideogram's 0.97 English OCR figure is a vendor benchmark. Sume serves Ideogram 4.5, not 4.0, so run 20 of your own headlines through OCR and count matches.

Ideogram's 4.0 technical post claims 0.97 X-Omni English OCR accuracy for a 9.3B-parameter open-weight model released June 3, 2026, and says text is controlled with JSON bounding boxes (Ideogram, read 2026-10-04). That is the vendor's own benchmark on its own test. Your headlines are not that test.
What Sume serves
Sume does not list Ideogram 4.0. The catalog has ideogram/ideogram-v4.5 and ideogram/ideogram-v3 (Sume Image API). Ideogram 4.5 takes quality low, medium or high, with list prices of $0.03, $0.06 and $0.22 per image, and resolution 1K or 2K. So the question for a buyer is not whether 0.97 is true, but how often your own copy comes out spelled right on the model you can call.
| Item | Ideogram 4.0 (vendor post) | Sume catalog |
|---|---|---|
| Model | 9.3B open weights, June 3, 2026 | ideogram-v4.5 and ideogram-v3 |
| Text claim | 0.97 X-Omni English OCR accuracy | No published score; test it |
| Quality tiers | Not applicable to the weights | low $0.03, medium $0.06, high $0.22 list |
A 20-headline exact-match test
Count a headline as a pass only at a ratio of 1.0, and look at the failures by eye because OCR can also be wrong. If a call returns 202, read the job result instead.
import difflib
import io
import os
import requests
import pytesseract
from PIL import Image
h = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def score(text: str) -> float:
body = {
"model": "ideogram/ideogram-v4.5",
"prompt": f'A poster with the headline "{text}" in bold type',
"quality": "medium",
}
r = requests.post("https://api.sume.com/v1/images", headers=h, json=body, timeout=60)
r.raise_for_status()
img = Image.open(io.BytesIO(requests.get(r.json()["data"][0]["url"], timeout=60).content))
read = pytesseract.image_to_string(img).lower().split()
return difflib.SequenceMatcher(None, " ".join(read), text.lower()).ratio()
print(score("Autumn sale starts Friday"))
Sources
Related posts
More in Models
- Music from a thumbnail: image_url on Sume's music router
Generate a music bed that matches a still: pass one public HTTPS image_url with the prompt to Sume's Music Router, and clear it with null when reusing objects.
- Image reference limits: GPT Image 2.5 takes 16, Ideogram 4.5 takes 5
GPT Image 2.5 accepts up to 16 input references on Sume; Ideogram 4.5 edits the first image and uses up to 4 more. How to order them and when to prefer which.
- Index-Translate 2B, 9B or 35B-A3B for subtitle translation
Index-Translate ships 2B, 9B and 35B-A3B text models. What the vendor pages say about each size, and how to pick one for subtitle cues you burn with Sume.
- Index-Translate, Echo, Homura: which part do you need?
Bilibili's Index-Translate release has text, speech, dubbing and length parts. Which fit subtitles, and which Sume step follows.
Written by Sume