Fix the text in a reference image before it goes into an H3 clip
Spend one cheap Ideogram 4.5 edit on a reference image's text before you spend on a MiniMax H3 clip that copies it. The order on Sume, with a request.

If a reference image has a wrong word on it, fix the image before you make the video. A video model that uses the picture as a reference carries its flaws into every frame, and a clip costs more than an image. On Sume that means one POST /v1/images Ideogram 4.5 edit first, then the corrected image's URL in the video request.
Why the order matters
Prices differ by orders of magnitude. The Image API docs list Fal list prices of $0.03, $0.06 and $0.22 per image for Ideogram 4.5 at low, medium and high. A video job is billed per output second at the provider list times 1.25, per the Video Router docs. A wrong word you find after rendering costs a clip; found before, it costs one low-tier image.
Ideogram's launch post (read 2026-10-05) says 4.5 is the most precise edit model and that it is available in its API. A third-party summary (the Hugging Face blog, read 2026-10-05) says MiniMax H3 takes up to 9 reference images. Together they suggest a two-step flow, which is what the docs let you build.
Where text breaks in video
Even with a perfect reference, small text in motion can change from frame to frame. A reference image sets the starting look, but the video model still draws every frame. That is why large, short words survive better than a paragraph on a label. For packaging, plan the shot so the label is large and steady in the frame for a moment.
If the text must be exact in the clip, the dependable plan is to keep it out of the generated footage and add it afterwards as captions or an overlay. Sume's docs list captions and timeline assembly as separate steps that work on Sume-hosted files.
Using several corrected images
The same order applies when you use many references. Fix the text on every image that has any, one edit each at the low tier, then submit the video with the corrected set. A third-party summary lists up to 9 images for H3, so a full set could be nine low-tier edits first. Count that against the clip cost before you decide: for a short draft clip, it may not be worth fixing text that will be too small to read anyway.
The two calls
Step 1 edits the image. Step 2 uses the Sume-hosted result URL as a reference.
# 1. Fix the word in the reference image
POST /v1/images
{"model": "ideogram/ideogram-v4.5", "quality": "low",
"prompt": "Change \"FRESHLY RAOSTED\" to \"FRESHLY ROASTED\". Keep everything else identical.",
"input_references": [{"type": "image_url",
"image_url": {"url": "https://example.com/bag.png"}}]}
# 2. Use data[0].url from the response in the video request
POST /v1/video-router/generate
{"model": "minimax-h3-max", "duration": 6, "resolution": "480p",
"prompt": "Slow push-in on the coffee bag from the reference image",
"reference_image_urls": ["<data[0].url from step 1>"]}A check between the steps
Read the corrected image at full size before you submit step 2. A completed edit is billed whether or not the word is right, per the Image API docs, so check the draft; then promote to a higher tier if the word came out right but the image is soft.
Draft the video at 480p. Do the final at 768p or 1080p once the reference is settled. If the clip still garbles small text, the usual fix is the same as for any generated video: keep text large and short, or add it later as captions instead of asking the model to render it.
A short example of the order of work makes it concrete. Take a product photo with a misspelled label. Edit the photo at the low tier and read the label. Repeat until the label is right. Only then submit the video request with the corrected image as a reference. The video job is the expensive step, so the cheap step goes first. Keep the scope of this advice in view. It rests on the Sume docs and the vendor pages named in the sources, read on 2026-10-05, and on nothing measured by Sume. Where a behavior depends on your own images, such as how a model redraws a certain typeface, run a small pilot at the low quality tier and judge the result yourself before you plan a batch. Write down the prompt, the model id and the quality tier you used, so the run can be repeated. When the catalog or the docs change, re-read them; the live catalog is the contract, and a post is only a snapshot of it.
Sources
Related posts
More in Media tools
- Godot MP3 loop: begin point only, so loop a Sume track as wav
Godot gives MP3 and Ogg only a loop begin point, no loop end. For a Sume BGM loop, cut a sample-exact wav with timeline audio split and set the loop on import.
- Graduation video music: a 2-minute anthem with a build
Brief a 2-minute graduation video bed with a build to a triumphant moment at 1:20 for the diploma reel: one generation and a render, $0.325 on Sume.
- Grayscale value check for AI product images: does it read at a glance?
Convert a Sume image to grayscale, measure the contrast between product and background with ImageStat, and flag low-value-contrast images before they ship.
- H3 Max Recast: 1-4 person photos, 768p default, 1080p opt-in
H3 Max Recast swaps the people in one 5-30 s video using 1-4 reference photos, one per new person. Output is 768p by default; 1080p is opt-in.
Written by Sume