Fix the text in a reference image before it goes into an H3 clip

Spend one cheap Ideogram 4.5 edit on a reference image's text before you spend on a MiniMax H3 clip that copies it. The order on Sume, with a request.

5 min readSume
All posts

If a reference image has a wrong word on it, fix the image before you make the video. A video model that uses the picture as a reference carries its flaws into every frame, and a clip costs more than an image. On Sume that means one POST /v1/images Ideogram 4.5 edit first, then the corrected image's URL in the video request.

Why the order matters

Prices differ by orders of magnitude. The Image API docs list Fal list prices of $0.03, $0.06 and $0.22 per image for Ideogram 4.5 at low, medium and high. A video job is billed per output second at the provider list times 1.25, per the Video Router docs. A wrong word you find after rendering costs a clip; found before, it costs one low-tier image.

Ideogram's launch post (read 2026-10-05) says 4.5 is the most precise edit model and that it is available in its API. A third-party summary (the Hugging Face blog, read 2026-10-05) says MiniMax H3 takes up to 9 reference images. Together they suggest a two-step flow, which is what the docs let you build.

Where text breaks in video

Even with a perfect reference, small text in motion can change from frame to frame. A reference image sets the starting look, but the video model still draws every frame. That is why large, short words survive better than a paragraph on a label. For packaging, plan the shot so the label is large and steady in the frame for a moment.

If the text must be exact in the clip, the dependable plan is to keep it out of the generated footage and add it afterwards as captions or an overlay. Sume's docs list captions and timeline assembly as separate steps that work on Sume-hosted files.

Using several corrected images

The same order applies when you use many references. Fix the text on every image that has any, one edit each at the low tier, then submit the video with the corrected set. A third-party summary lists up to 9 images for H3, so a full set could be nine low-tier edits first. Count that against the clip cost before you decide: for a short draft clip, it may not be worth fixing text that will be too small to read anyway.

The two calls

Step 1 edits the image. Step 2 uses the Sume-hosted result URL as a reference.

# 1. Fix the word in the reference image
POST /v1/images
{"model": "ideogram/ideogram-v4.5", "quality": "low",
 "prompt": "Change \"FRESHLY RAOSTED\" to \"FRESHLY ROASTED\". Keep everything else identical.",
 "input_references": [{"type": "image_url",
   "image_url": {"url": "https://example.com/bag.png"}}]}

# 2. Use data[0].url from the response in the video request
POST /v1/video-router/generate
{"model": "minimax-h3-max", "duration": 6, "resolution": "480p",
 "prompt": "Slow push-in on the coffee bag from the reference image",
 "reference_image_urls": ["<data[0].url from step 1>"]}

A check between the steps

Read the corrected image at full size before you submit step 2. A completed edit is billed whether or not the word is right, per the Image API docs, so check the draft; then promote to a higher tier if the word came out right but the image is soft.

Draft the video at 480p. Do the final at 768p or 1080p once the reference is settled. If the clip still garbles small text, the usual fix is the same as for any generated video: keep text large and short, or add it later as captions instead of asking the model to render it.

A short example of the order of work makes it concrete. Take a product photo with a misspelled label. Edit the photo at the low tier and read the label. Repeat until the label is right. Only then submit the video request with the corrected image as a reference. The video job is the expensive step, so the cheap step goes first. Keep the scope of this advice in view. It rests on the Sume docs and the vendor pages named in the sources, read on 2026-10-05, and on nothing measured by Sume. Where a behavior depends on your own images, such as how a model redraws a certain typeface, run a small pilot at the low quality tier and judge the result yourself before you plan a batch. Write down the prompt, the model id and the quality tier you used, so the run can be repeated. When the catalog or the docs change, re-read them; the live catalog is the contract, and a post is only a snapshot of it.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume