Talking Still Testimonial: VEED Fabric 1.0 or H3 Max Lip Sync, Priced

Make a 14.8-second talking-head clip from a still and a TTS line on Sume: Fabric at 720p versus MiniMax H3 Max Lip Sync at 768p, with each price multiplied out.

4 min readSume
All posts

The short answer

For a 14.8-second talking clip from one still and one audio line, MiniMax H3 Max Lip Sync at 768p costs $1.480 and VEED Fabric 1.0 at 720p costs $2.775 on Sume list prices, plus a few cents for the speech. Both take a still and an audio file. Sume's docs say a person who speaks on camera is Fabric with an accepted still and TTS, because video models do not lip-sync to generated TTS or to a later voice-over.

What the docs say about the two routes

The models page lists VEED Fabric 1.0 (POST /v1/veed/fabric-1.0) as a talking still plus audio clip, and MiniMax H3 Max Lip Sync (POST /v1/minimax/h3-max/lip-sync) with the same still-plus-audio body, audio 5 to 14.8 seconds, billed at list times 1.25. For Fabric, send audio_url, a measured duration_seconds and only one visual source: image_url of an inspected posed still, or an avatar_handle when the user named that avatar.

Rates (Sume catalog, read 2026-10-08)
ModelRate14.8 s clip5 s clip
VEED Fabric 1.0, 720p$0.1875 per audio second$2.775$0.9375
VEED Fabric 1.0, 480p$0.10 per audio second$1.480$0.500
MiniMax H3 Max Lip Sync, 768p$0.10 per audio second$1.480$0.500
MiniMax H3 Max Lip Sync, 480p$0.0625 per audio second$0.9250$0.3125

The chain

Write a one-breath line, generate it with tts_create ($0.0475 per 1,000 characters, so a 250-character line is $0.0119), measure its duration, then submit the still and audio. Measure the length from the audio file, not the script, because the price is per audio second and Fabric needs the measured value.

Use a still where the face is clear, forward-facing and not covered by the hand or product. Inspect the still with an image check before you spend on the talking pass. If the product must be in frame, keep it away from the mouth.

Which to pick comes down to the audio length and look. H3 Max Lip Sync accepts audio of 5 to 14.8 seconds, so it fits one short line; Fabric is the documented route for the longer talk. Make one test clip at the lower-cost 480p setting on each route, compare the lip movement on the same still, and only then pay for the 720p or 768p render. Both tests together at 480p and a 5-second line are $0.8125.

  • Only use a face and a testimonial text you have the right to use. A synthetic testimonial that presents invented customer experience as real can break advertising rules; label it as a dramatization.
  • For lines over 14.8 seconds, split the audio and join the clips in Timeline 1.0.
  • Re-run only the line that failed: 14.8 seconds at Fabric 720p is $2.775 per retry.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume