Make a Talking Photo Speak Spanish: TTS Language, Then Fabric

Two steps on Sume: generate Spanish speech with the language set, then animate a still with Fabric. Costs for a 30-second clip and what Avatar Video can't do.

5 min readSume
All posts

To make a still photo speak Spanish on Sume, generate the Spanish audio first with the text-to-speech language set to es, then send that audio and the still to Fabric. Avatar Video itself should not be used for Spanish: Sume's guidance is that scripts there should be English, because non-English speech breaks caption alignment.

Prices are from API pricing, read on 2026-10-05: text to speech is $0.0475 per 1,000 characters, and Fabric is $0.1875 per second at 720p.

What are the two steps?

Step one, text to speech: send the Spanish transcript, a voice and language. If the voice's language and language disagree, the API can refuse the request with a 409 tts_voice_language_mismatch before any charge, and you can only continue by setting confirm_language_mismatch. Pick a Spanish voice instead of overriding.

Step two, Fabric: send the generated audio as audio_url (it must be Sume-hosted, 10 MB at most), the measured duration_seconds, and one visual source. The preflight checks only host and size, not the language, so a mismatched voice is yours to catch by listening.

POST https://api.sume.com/v1/veed/fabric-1.0
{
  "image_url": "https://media.sume.com/portrait.png",
  "audio_url": "https://media.sume.com/saludo-es.mp3",
  "duration_seconds": 30,
  "resolution": "720p"
}

What does a 30-second clip cost?

About 75 Spanish words is roughly 450 characters, which is $0.02 of speech at $0.0475 per 1,000 characters. The Fabric part dominates: 30 seconds at 720p is $5.63 (30 x $0.1875 = $5.625), or $3.00 at 480p.

30-second Spanish talking photo (rates read 2026-10-05; character count assumed)
StepRateAmount
Text to speech, about 450 characters$0.0475 per 1,000 characters$0.02
Fabric at 720p, 30 s$0.1875 per second$5.63
Fabric at 480p, 30 s$0.10 per second$3.00

What does this not do?

It does not translate: you supply the Spanish transcript. It does not converse; the still moves to the audio you gave it. For Avatar Video scenes in other languages, see the limits in Generate avatar video and keep spoken scripts English.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume