LCP hero image from an AI API: self-host, size it, fetchpriority
web.dev treats an LCP of 2.5 s or less as good. Copy a Sume hero image to your own host, set width and height, and use fetchpriority high; do not lazy-load it.

web.dev's LCP guidance says sites should aim for an LCP of 2.5 seconds or less for at least 75% of page visits. If your hero image comes from an image API, the three things in your control are where the bytes are served from, whether the browser can size the box before the image arrives, and the priority you give the fetch.
Where the time goes
web.dev splits LCP into four subparts and gives guideline shares. The two delays should be small, and most of the time should be spent on the server's first byte and on actually downloading the image.
| Measure | Guidance |
|---|---|
| Good LCP | 2.5 s or less at the 75th percentile of page visits |
| Time to first byte | About 40% |
| Resource load delay | Under 10% |
| Resource load duration | About 40% |
| Element render delay | Under 10% |
Self-host the Sume result
The Image API returns data[].url, a Sume-hosted media.sume.com URL. A page that must load fast for every visitor should not depend on a third-party origin you do not control. Download the bytes once at build or publish time and serve them from your own origin or CDN, where you control caching and the first-byte time. Ask for output_format: "webp" or jpeg for a photographic hero; the format list is png, jpeg, webp and svg.
Markup for the hero
MDN says that including width and height lets the browser calculate the aspect ratio before the image loads, which prevents layout shift. MDN also cautions that lazy-loaded images in the visual viewport may not be visible when the window load event fires, so leave loading="lazy" off a hero that is above the fold. web.dev reports that setting fetchpriority="high" improved one LCP from 2.6 s to 1.9 s, and that it matters less when you have already preloaded the image early in the head.
<img
src="/images/hero-1600.webp"
srcset="/images/hero-800.webp 800w, /images/hero-1600.webp 1600w"
sizes="100vw"
width="1600" height="900"
fetchpriority="high"
alt="Ceramic mug of coffee on a wooden table">Generate to the layout
A 16:9 hero maps to aspect_ratio: "16:9", which most catalog models list. Generate once, save two widths, and write the width and height attributes from the real file rather than from the request, since the pixels a model returns depend on its tier.
import os, io, requests
from PIL import Image
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={"model": "bytedance-seed/seedream-4.5", "aspect_ratio": "16:9", "prompt": "wide product scene on a clean surface, soft light, no text"},
timeout=90,
)
r.raise_for_status()
if r.status_code != 200:
raise SystemExit("202: read the finished job from /v1/jobs/{id}/result")
img = Image.open(io.BytesIO(requests.get(r.json()["data"][0]["url"], timeout=60).content))
img = img.convert("RGB")
for w in (800, 1600):
out = img.resize((w, round(w * img.height / img.width)), Image.LANCZOS)
out.save(f"hero-{w}.webp", "WEBP", quality=82)
print(f"hero-{w}.webp", out.size)A short measurement habit
Change one thing at a time and compare the four subparts before and after. If the resource load delay is large, the browser found the image late: move it from a CSS background to an img element, or preload it. If the resource load duration is large, the file is too heavy: lower the quality or width of the WebP. If the render delay is large, something else is blocking paint. Generating a better picture does not move any of these numbers; shrinking and serving it correctly does.
What not to expect
A fast image does not fix a slow server: web.dev's guideline gives the first byte about 40% of the budget. Measure the four subparts in your own field data before changing the image pipeline.
The target is stated for page visits, so field data is the measure and a lab run on one machine is only a guide.
Sources
Related posts
More in Developers
- Listen to This Article: Audio Player With Read-Along Highlighting
Build a read-along article player: ask Sume TTS for word timestamps and sentence segments, save the timings as JSON, and highlight text from currentTime.
- LiteLLM /mcp-rest/tools/call: test Sume tools without an LLM
LiteLLM exposes REST routes to list and call MCP tools with no model in the loop. Use them to smoke-test Sume's read tools and gates before an agent sees them.
- Live AI avatar API: a Tavus conversation vs a Sume job
A live avatar API creates a room you join. Sume's Avatar API creates a job you poll. Field-by-field map of Tavus create conversation and Sume talking-video.
- LM Studio mcp.json: add Sume's hosted MCP server with an API key
LM Studio 0.3.17 and later accept remote MCP servers in mcp.json. The exact entry for https://mcp.sume.com/mcp with a Bearer header, plus the gates to know.
Written by Sume