GPT Image 2.5 portrait into a Fabric talking clip: 10 s costs $1.89

A ChatGPT Image 2.5 portrait plus a 10-second VEED Fabric 1.0 talking clip costs about $1.89 on Sume at 720p. What to send and what Fabric needs.

5 min readSume
All posts

A talking-head clip made from a ChatGPT Image 2.5 portrait and 10 seconds of audio costs about $1.89 on Sume: $0.0165 for a medium 1024x1024 still and $1.875 for the Fabric clip at $0.1875 per second at 720p. The still is a repository estimate and the Fabric rate is from the models page.

Fabric is veed/fabric-1.0, called at POST /v1/veed/fabric-1.0, and it turns one still plus one audio file into a talking clip.

What the request needs

The models page says to send audio_url, a measured duration_seconds, and only one visual source. The preferred source is the image_url of the generated, inspected posed still. avatar_handle is for when the user named an avatar, and you cannot send both.

The audio_url must be a public HTTPS URL on the Sume media host (typically a TTS segment, max 10 MB); other hosts are rejected. image_url must be a public HTTPS still. Duration is billed by the second, so measure the audio rather than guessing.

{
  "image_url": "https://media.sume.com/img/EXAMPLE/0.png",
  "audio_url": "https://media.sume.com/artifacts/artf_demo/voice.mp3",
  "duration_seconds": 10
}

Cost by length

One still, then Fabric at $0.1875 per second at 720p.

Portrait still plus Fabric talking clip at 720p (read 2026-10-07)
Audio lengthFabric clipStillTotal
5 s$0.94$0.0165$0.95
10 s$1.88$0.0165$1.89
15 s$2.81$0.0165$2.83
30 s$5.63$0.0165$5.64

Planning the audio first

Because Fabric bills by the second of audio, write and record the line before you generate anything visual. Measure the real file length and pass it as duration_seconds, since a padded number raises the credits reserved.

Text-to-speech output is a convenient source for the audio file, because it is already on the Sume media host. The file should contain only the speech, with no long silences at the head or tail that you would pay for.

  • Trim leading and trailing silence before upload.
  • Keep one speaker per clip.
  • duration_seconds accepts 1 to 300; for a longer scene, split the audio and join the clips in a timeline.

Make the still usable

Fabric animates the face you give it, so the still should be a front-facing portrait with a closed mouth, plain background and no hands near the face. Generate three options with n, inspect them, and pass only the best one. Do not run Fabric on a still you have not looked at, since a bad portrait bills the same as a good one.

The older POST /v1/avatar-1.0/image-to-video route still works as an alias, and the fabric experimental route is test-only. Use veed/fabric-1.0 for anything you build on.

For a presenter who appears in many videos, keep one approved portrait and reuse it. Each new clip then costs the Fabric seconds only, with no new image bill. At $0.1875 per second, a minute of talking head is $11.25 as separate clips.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume