Talking photo AI without watermark: what Sume documents

Sume's docs describe no watermark option for talking photo clips. They do fix the inputs: a still, Sume-hosted audio under 10 MB, 1-300 s or 5-14.8 s by route.

4 min readSume
All posts

Sume's docs say nothing about a watermark on talking photo output, either way, so this post makes no claim about one. What they do document is the request: a still image plus audio hosted on Sume's media host, sent to one of two routes that differ in how long the audio may be. Check a first result yourself before you build on it.

The limits below come from Sume's OpenAPI reference and Video models page, read 2026-09-29. Nothing here describes any other vendor's free tools or their watermark policy.

Which routes turn a photo and audio into a talking clip?

Both routes take the same body: one visual source, either image_url or a ready avatar_id or avatar_handle, plus audio_url. They differ in the audio window and the resolutions offered.

Talking-photo routes on Sume, read 2026-09-29.
RouteAudio lengthResolutions
POST /v1/veed/fabric-1.01 to 300 seconds480p, 720p (default 720p)
POST /v1/minimax/h3-max/lip-sync5 to 14.8 seconds480p, 768p, 1080p (default 768p)

What rules does the audio file have to meet?

It must be a public HTTPS URL on the Sume media host. The schema says non-Sume hosts are rejected, and the file may be at most 10 MB. In practice the audio is often a text-to-speech segment made through Sume first. Pass duration_seconds as well: it sets the credit reservation, up to 300 on the first route.

What happens with audio outside 5 to 14.8 seconds?

On the second route, the request is refused. The schema explains that the provider rejects shorter audio and silently clips longer audio, so Sume refuses requests outside the window instead of clamping them. Output length follows the audio. If your line runs longer, use the first route or split it, as covered in lip sync for a clip longer than 15 seconds.

How do I send a still and audio?

Use image_url for a public HTTPS still and audio_url from Sume's media host.

curl -X POST https://api.sume.com/v1/veed/fabric-1.0 \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image_url": "https://example.com/portrait.png",
    "audio_url": "https://media.sume.com/example/line.wav",
    "duration_seconds": 9,
    "resolution": "720p"
  }'

Which route should I pick for a short line?

If your audio fits 5 to 14.8 seconds, both routes accept it, and the choice comes down to the resolutions you want: the first route stops at 720p, while the second offers 768p and 1080p. Sume's models page describes MiniMax H3 Max as the faster 768p variant. If the line is shorter than 5 seconds or longer than 14.8, only the first route applies, up to 300 seconds. Because the first route's credit reservation is based on duration_seconds, send an honest value rather than the maximum, or you will reserve more than the clip needs.

Send the still as a clear, front-facing portrait with the mouth visible. The docs do not set image rules beyond a public HTTPS URL, so keep that URL fetchable and unsigned where you can.

How do I find out about a watermark?

Render one short clip at your intended resolution and inspect the finished video before you scale up. If output details matter to your use, treat what you see as the truth, and ask Sume support before relying on anything the docs do not state. For the input side in more depth, read Lip sync API: photo and audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume