Make a Talking Photo Speak Spanish: TTS Language, Then Fabric
Two steps on Sume: generate Spanish speech with the language set, then animate a still with Fabric. Costs for a 30-second clip and what Avatar Video can't do.

To make a still photo speak Spanish on Sume, generate the Spanish audio first with the text-to-speech language set to es, then send that audio and the still to Fabric. Avatar Video itself should not be used for Spanish: Sume's guidance is that scripts there should be English, because non-English speech breaks caption alignment.
Prices are from API pricing, read on 2026-10-05: text to speech is $0.0475 per 1,000 characters, and Fabric is $0.1875 per second at 720p.
What are the two steps?
Step one, text to speech: send the Spanish transcript, a voice and language. If the voice's language and language disagree, the API can refuse the request with a 409 tts_voice_language_mismatch before any charge, and you can only continue by setting confirm_language_mismatch. Pick a Spanish voice instead of overriding.
Step two, Fabric: send the generated audio as audio_url (it must be Sume-hosted, 10 MB at most), the measured duration_seconds, and one visual source. The preflight checks only host and size, not the language, so a mismatched voice is yours to catch by listening.
POST https://api.sume.com/v1/veed/fabric-1.0
{
"image_url": "https://media.sume.com/portrait.png",
"audio_url": "https://media.sume.com/saludo-es.mp3",
"duration_seconds": 30,
"resolution": "720p"
}What does a 30-second clip cost?
About 75 Spanish words is roughly 450 characters, which is $0.02 of speech at $0.0475 per 1,000 characters. The Fabric part dominates: 30 seconds at 720p is $5.63 (30 x $0.1875 = $5.625), or $3.00 at 480p.
| Step | Rate | Amount |
|---|---|---|
| Text to speech, about 450 characters | $0.0475 per 1,000 characters | $0.02 |
| Fabric at 720p, 30 s | $0.1875 per second | $5.63 |
| Fabric at 480p, 30 s | $0.10 per second | $3.00 |
What does this not do?
It does not translate: you supply the Spanish transcript. It does not converse; the still moves to the audio you gave it. For Avatar Video scenes in other languages, see the limits in Generate avatar video and keep spoken scripts English.
Sources
Related posts
More in Developers
- Map a bulk queue item index back to a SKU: keep your own ledger
Queue items return index, status and run_id, not your SKU. Save a SKU-by-index table at submit time and join it to the queue receipt when you poll.
- Marketing API v24.0 ends Oct 6, 2026: pin the version in your uploader
Meta lists Marketing API v24.0 as available until October 6, 2026 and v25.0 as latest. Make the version a setting, and keep Sume render jobs separate.
- Marketing API v24 expiry runbook: replay the upload, not the render
When Meta's v24.0 window closes on October 6, 2026, fix uploads by replaying stored Sume job results. The same Idempotency-Key never bills a render twice.
- Match STT results to file ids: the job echoes your Idempotency-Key
Sending 300 recordings to Sume STT? Set Idempotency-Key to your own file id. The job record returns it, so results map back after a crash or a retry.
Written by Sume