Fetch tool to media input: which URLs Sume accepts
A model can fetch a page, but Sume media inputs must be public HTTPS image or video URLs, not web pages. Here is what each Sume endpoint accepts.

A web fetch tool returns page content, but Sume's generation inputs must be fetchable public HTTPS URLs of the media file itself: an image or a video, with a matching content type. Have your agent extract the direct media URL from the page before it calls Sume, and do not pass the page URL.
What Vercel announced
The Vercel changelog for Sep 30, 2026 says AI Gateway now supports Browserbase Search and Fetch, giving models web discovery. That helps an agent find reference pages. It says nothing about media inputs, which is where Sume's rules apply.
What Sume accepts
Launch generation requests take media as public HTTPS URLs in the fields named by the OpenAPI schema. Localhost, private-network, non-HTTPS, signed or private URLs, and mismatched content types are rejected before generation is submitted.
| Workflow | Field |
|---|---|
| Avatar from a photo | input.image_url |
| Avatar video product image | product_image |
| Avatar video scene photo | scene.image_url |
| Face swap (Beta) source video | video_url |
| Video captions | video_url |
Sume-hosted sources for editing tools
Some tools are stricter. Video trim and Timeline 1.0 take only this workspace's media.sume.com artifacts or assets, and reject off-host URLs with unsupported_media_source. Import first with POST /v1/media-imports (media-imports_create over MCP), then use the resulting Sume URL.
A guard for your agent
Before calling Sume, check that the candidate is an HTTPS URL that ends in a file the agent can identify as media. This small Python check catches the common mistake of passing a page URL; the server still does the final validation.
from urllib.parse import urlparse
import requests
def looks_like_media(url: str) -> bool:
u = urlparse(url)
if u.scheme != "https" or not u.hostname:
return False
r = requests.head(url, allow_redirects=True, timeout=15)
kind = r.headers.get("content-type", "").split(";")[0]
return r.ok and kind.startswith(("image/", "video/"))Trending videos is a different case: it returns public watch URLs for research and does not mirror downloadable source files, so those URLs are not generation inputs. See Media inputs.
Sources
Related posts
More in Developers
- What a media MCP server should declare at server/discover
MCP 2026-07-28 adds a required server/discover call. A media server has more to say than versions: async jobs, wait limits, scopes. Where Sume documents each.
- Which Lyria ran? job.model vs job.request.routed_model on Sume
On Sume's Music Router, job.model echoes what you sent and job.request.routed_model names the engine that ran. How to read both and pin a Lyria id.
- whisper-1 or gpt-transcribe for subtitles: what OpenAI assigns to each
OpenAI recommends gpt-transcribe, gpt-4o-transcribe-diarize for speakers, whisper-1 for translation and subtitles. Plus the 25 MB limit and a chunking script.
- Why adaptive and auto are missing from /v1/videos/models
Seedance and Wan list auto and adaptive aspect ratios in the Video Router catalog, but /v1/videos/models filters them out. Which models, and a script to see it.
Written by Sume