GPT Image 2.5 portrait into a Fabric talking clip: 10 s costs $1.89
A ChatGPT Image 2.5 portrait plus a 10-second VEED Fabric 1.0 talking clip costs about $1.89 on Sume at 720p. What to send and what Fabric needs.

A talking-head clip made from a ChatGPT Image 2.5 portrait and 10 seconds of audio costs about $1.89 on Sume: $0.0165 for a medium 1024x1024 still and $1.875 for the Fabric clip at $0.1875 per second at 720p. The still is a repository estimate and the Fabric rate is from the models page.
Fabric is veed/fabric-1.0, called at POST /v1/veed/fabric-1.0, and it turns one still plus one audio file into a talking clip.
What the request needs
The models page says to send audio_url, a measured duration_seconds, and only one visual source. The preferred source is the image_url of the generated, inspected posed still. avatar_handle is for when the user named an avatar, and you cannot send both.
The audio_url must be a public HTTPS URL on the Sume media host (typically a TTS segment, max 10 MB); other hosts are rejected. image_url must be a public HTTPS still. Duration is billed by the second, so measure the audio rather than guessing.
{
"image_url": "https://media.sume.com/img/EXAMPLE/0.png",
"audio_url": "https://media.sume.com/artifacts/artf_demo/voice.mp3",
"duration_seconds": 10
}Cost by length
One still, then Fabric at $0.1875 per second at 720p.
| Audio length | Fabric clip | Still | Total |
|---|---|---|---|
| 5 s | $0.94 | $0.0165 | $0.95 |
| 10 s | $1.88 | $0.0165 | $1.89 |
| 15 s | $2.81 | $0.0165 | $2.83 |
| 30 s | $5.63 | $0.0165 | $5.64 |
Planning the audio first
Because Fabric bills by the second of audio, write and record the line before you generate anything visual. Measure the real file length and pass it as duration_seconds, since a padded number raises the credits reserved.
Text-to-speech output is a convenient source for the audio file, because it is already on the Sume media host. The file should contain only the speech, with no long silences at the head or tail that you would pay for.
- Trim leading and trailing silence before upload.
- Keep one speaker per clip.
duration_secondsaccepts 1 to 300; for a longer scene, split the audio and join the clips in a timeline.
Make the still usable
Fabric animates the face you give it, so the still should be a front-facing portrait with a closed mouth, plain background and no hands near the face. Generate three options with n, inspect them, and pass only the best one. Do not run Fabric on a still you have not looked at, since a bad portrait bills the same as a good one.
The older POST /v1/avatar-1.0/image-to-video route still works as an alias, and the fabric experimental route is test-only. Use veed/fabric-1.0 for anything you build on.
For a presenter who appears in many videos, keep one approved portrait and reuse it. Each new clip then costs the Fabric seconds only, with no new image bill. At $0.1875 per second, a minute of talking head is $11.25 as separate clips.
Sources
Related posts
More in Sume Avatar 1.0
- Griffin-style follow-ups with rendered avatar clips and branching
Griffin-Lite reacts live. Until you can use it, approximate a guided conversation with a set of pre-rendered Sume avatar clips and your own branching logic.
- Interactive avatar or avatar video? A five-question test
Griffin-Lite is a research preview. Five questions tell you whether you need a live avatar or a rendered Sume Avatar 1.0 clip, and what to ship this quarter.
- Name your avatars: a handle scheme that fits 2 to 30 characters
Sume avatar handles allow letters, digits, periods and underscores, 2 to 30 characters. Build a team naming scheme that passes validation and stays readable.
- Let users pick an avatar in your app with GET /v1/avatar-1.0/avatars
Build an avatar picker on the Sume Avatar 1.0 list route: read ready avatars server-side, cache them, and pass the chosen handle to talking-video.
Written by Sume