AI caricature from a photo via API: reference edit on Sume
Turn your own photo into a caricature with a reference edit: which Sume rows take a photo, what a call costs from $0.025, and why n=4 helps. Python example.

A caricature from a photo on the Sume image API is a reference edit: you pass the photo as an input_references entry and ask for an exaggerated illustrated version of the same person. Sixteen of the catalog rows take references, and the cheapest are Qwen Image and Grok Imagine at $0.025 per image, read 2026-10-03. Soul, Qwen Image Max, Imagen 4 Fast, Imagen 4 Ultra and Recraft V4 are text-to-image only and reject a photo.
Use photos of yourself or of people who agreed to it. The photo has to be at a public HTTPS URL, because localhost, private-network and non-HTTPS URLs are rejected before submission. Sume does not host your upload for you in this call, so put the file at a URL you control.
Rows worth trying
| Model | Catalog id | Per image | References |
|---|---|---|---|
| Grok Imagine | x-ai/grok-image | $0.025 | 10 |
| Qwen Image | qwen/qwen-image | $0.025 | 10 |
| Seedream 4.0 | bytedance-seed/seedream-4 | $0.0325 | 10 |
| Flux 2 Pro | black-forest-labs/flux.2-pro | $0.0375 | 10 |
| Seedream 5.0 Lite | bytedance-seed/seedream-5-lite | $0.04375 | 10 |
| Seedream 4.5 | bytedance-seed/seedream-4.5 | $0.05 | 10 |
The request
Describe the style and what to exaggerate, and say what to keep (the hair, glasses, the general likeness). On an edit, prefer aspect_ratio: "auto" where the model lists it so the output matches the photo's shape; Grok Imagine does not list auto, so give it an explicit ratio. Grok Imagine returns one image per call. A call that is still running after the 30-second wait returns 202 with a job envelope instead of 200, so check the status code before reading data; see Jobs and results.
import os
import requests
resp = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={
"model": "bytedance-seed/seedream-4",
"prompt": "playful caricature of the person in the photo, big head, small body, "
"keep the glasses and hairstyle, bold ink lines, flat colors",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/me.jpg"}}
],
"aspect_ratio": "auto",
"n": 4,
},
timeout=60,
)
resp.raise_for_status()
print(resp.status_code, resp.json()["data"])Why generate four
Likeness varies from sample to sample, and Sume's v1 image API takes no seed (it returns 400 unsupported_parameter), so you cannot repeat a good one. Asking for n: 4 costs four times the price, for example $0.13 on Seedream 4.0, and gives you a pick. If none of the four works, change the prompt, not the count.
Limits to keep in mind
- A moderated or refused generation is a failure and is not billed, but retrying an unchanged prompt will refuse again.
- Seedream 4.5 and 5.0 Lite do not list
auto; the other rows differ, so readsupported_parameters. - Do not use another person's photo without their permission.
Sources
Related posts
More in Use cases
- AI certificate background via API: text-free art, names added in code
Generate a landscape certificate border at 4:3 or 3:2, with no text, then print names and dates with code so nothing is misspelled. Sume rows and cost.
- AI character series on Shorts: avoid the same situation each time
YouTube's inauthentic content policy flags characters in identical situations with the same outcomes. How to keep an AI character and vary the story in Sume.
- AI children's book illustrations with the same character on every page
Make storybook pages where the hero stays recognisable: draw one anchor image, pass it as a reference on every page call, and keep the style words fixed.
- AI classroom background music for lesson videos, under narration
Make calm instrumental music for a lesson video: a prompt that keeps vocals out, a Python script, and a Timeline bed that ducks under the teacher's voice.
Written by Sume