Talking presenter: Fabric still plus voice, or video model plus voice?

Sume docs: video models do not lip-sync to TTS, so use a still plus audio on Fabric ($0.1875/s at 720p listed). Wan 3.0 is cheaper but will not match the voice.

5 min readSume
All posts

For a person who speaks on camera, use a still plus an audio clip on VEED Fabric 1.0, not a video model with narration laid underneath. The Sume docs state it directly: video models do not lip-sync to generated TTS or to a later voice-over, so a face that talks is never a video-model clip with narration under it. Video models are for wordless beats, B-roll and product motion.

The routes that exist

These rows are from the Sume Models page, read 2026-10-08.

Talking-clip and video routes on Sume, read 2026-10-08
RouteInputWhat it is for
veed/fabric-1.0 (POST /v1/veed/fabric-1.0)image_url or avatar_handle (not both), audio_url, measured duration_secondsA talking still; the supported path for talking clips
minimax/h3-max/lip-syncSame still + audio body as Fabric; audio 5 to 14.8 sLip sync, billed at list x 1.25
Avatar videos (POST /v1/avatar-1.0/talking-video)A ready avatar handle and a script or scenes; 4 to 60 sScript-driven talking video
Video generation (POST /v1/videos)Prompt, frames, referencesWordless beats, B-roll, product motion

The price comparison that misleads

Fabric's legacy image-to-video route is listed in the docs at $0.1875 per second at 720p. A 15-second talking clip is therefore about 15 x $0.1875 = $2.81 at that listed rate. A Wan 3.0 clip of 15 s at 720p is $1.88, which is lower, but it does not give you a mouth that follows your voice track. Compare them as different products, not as two prices for the same result.

Order of work for a presenter ad

  • Write the script and generate the voice. Keep the audio file; Fabric takes a public HTTPS audio_url and a measured duration_seconds.
  • Generate and inspect a posed still. The docs prefer the generated, inspected still as the image_url; use avatar_handle only when the user named that avatar.
  • Send the still and audio to Fabric. Only one visual source is allowed per request.
  • Make B-roll and product shots with Auto or a pinned video model, and cut them together on the timeline.

When a video model is fine

If the person does not speak in the shot, such as a walk-in, a hand holding a product, or a reaction with a voice-over that is clearly off-screen, a video model is the right tool and costs less per second. Read the public-URL rules on Media inputs before you send audio or stills; Sume rejects localhost, private-network and non-HTTPS URLs before submission.

Check it before you run it

Every figure above is a catalog list price times 1.25, rounded up to the cent, as of 2026-10-08. Catalogs change, so before a large batch, read the current model entry in the docs and recompute the one line that matters for your case. Write the arithmetic next to the job in your own notes: list rate, seconds or characters, multiplier, rounding. If the result differs from the wallet charge by more than a cent, the catalog entry has changed, and the docs page is the place to find out why.

Run one small job first. Submit a single request with the pinned model id and the settings in the tables, poll the returned polling_url until it finishes, and compare the charge with your estimate. Then scale up. Using a pinned id for the test matters, because sume/auto never names the family, so you cannot tie its charge to the row you priced.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume