Arabic text to speech API: send language ar to Sume TTS
Arabic is on Cartesia's Sonic 3.6, 3.5 and 3 but never on Sonic 2. How to send ar through Sume TTS, which voice id to use, and what an Arabic script costs.

Yes: Sume TTS can speak Arabic. Send the transcript in Arabic script with language: "ar" and an Arabic-capable voice id to POST /v1/tts-1.0/generate. Sume passes the language through to the Cartesia engine, and Cartesia lists ar on Sonic 3.6, Sonic 3.5 and Sonic 3. Leave language out and the provider default is English, so Arabic text without the field is read with the wrong voice rules.
Three details decide whether it works first time: which model generation you are on, where the voice id comes from, and what Sume's own voice library does and does not cover.
Which Sonic generations speak Arabic?
Cartesia's older-models page gives each snapshot's language list. The Sonic 2 snapshots stop at English, French, German, Spanish, Portuguese, Chinese, Japanese and Korean, so a pipeline that was built on Sonic 2 never had Arabic. Cartesia also states that sonic-2 and sonic-turbo stop working after October 20, 2026. Sume's TTS 1.0 always uses the current stable Sonic (sonic-3.6), and the TTS Router offers sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview, so none of the Router ids is affected by that shutdown.
| Language | Code | sonic-3.6-2026-08-27 | sonic-3.5-2026-05-04 | sonic-3-2026-01-12 | In Sume's voice-library language list |
|---|---|---|---|---|---|
| Arabic | ar | Yes | Yes | Yes | No |
| Hebrew | he | Yes | Yes | Yes | No |
| English | en | Yes | Yes | Yes | Yes |
Where does the Arabic voice id come from?
The Sume voice selector takes a TTS voice UUID or a voi_ library id. A UUID is a Cartesia voice id, so you can pick an Arabic voice in Cartesia's own library and pass its UUID. Sume's tool guidance is explicit that Assets and Voices rows are optional references and not an admission gate: a raw voice id goes straight to the request.
The part that surprises people is the voice-creation side. Sume's voice tool (voices_create) accepts 16 languages: en, ko, ja, zh, es, fr, de, pt, it, hi, nl, pl, ru, sv, tr and tl. Arabic is not among them, so you cannot save a new Arabic voice into the Sume library, but you can still synthesize Arabic with a UUID you already hold.
What does a request look like, and what does it cost?
Set VOICE_ID to the Arabic voice UUID and SUME_API_KEY to your key. The call queues an async job and prints the job envelope; poll the job's status and fetch its result once it completes.
- Price: $0.0475 per 1,000 transcript characters, spaces and punctuation included, 20,000 characters maximum per request. A 4,000-character Arabic script is $0.19.
- Output defaults to MP3 at 44,100 Hz and 128 kbps; ask for
wavif you will join takes later. - Speed accepts 0.6 to 1.5 in
generation_config.speed, volume 0.5 to 2.
import os
import uuid
import requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={
"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"transcript": "مرحباً بكم في نشرتنا الأسبوعية.",
"language": "ar",
"voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
"mode": "async",
},
timeout=30,
)
r.raise_for_status()
print(r.json())What if the voice is not Arabic?
If the voice you pass is saved in Sume with a different primary language than the one you request, Sume answers a known mismatch with HTTP 409 tts_voice_language_mismatch before any job or charge, and the MCP tool turns it into a warning that needs the user's confirmation. A raw UUID Sume has no metadata for is not blocked: unknown library data never stops a voice id. That makes the audio check yours: listen to the first take before you queue a long script.
Should I worry about the Sonic 2 shutdown?
Only if an old pipeline pins it. Cartesia's older-models page marks sonic-2 and sonic-turbo as stopping on October 20, 2026, and the dated 2025 snapshot of Sonic 3 the same day. Sume's TTS 1.0 does not take a model id, and the Router ids listed above are newer, so a Sume request does not need migration. If you also call the provider directly elsewhere, move those calls to a current Sonic before the date.
Sources
Related posts
More in Models
- AudioCraft weights are CC-BY-NC: which models that covers
AudioCraft's code is MIT but its README puts the model weights under CC-BY-NC 4.0, including MusicGen and AudioGen. What that means for ad and video audio.
- AuK: MIT speech model that edits audio, and what Sume covers
AuK is Tencent's 1.5B MIT-licensed speech model for generation and editing. What its card lists, what it omits, and which parts Sume's audio tools cover.
- 21:9 architecture photos with AI: which Sume models take the ratio
Nano Banana 2 and Pro, FLUX.2 pro and flex, Qwen and Recraft list 21:9 on Sume. GPT Image 2.5 gets 3840x1648 by custom size. Imagen, Ideogram, Seedream do not.
- Can you sell images from open-weights models? Licences compared
Open weights do not mean commercial use. Ideogram 4, Qwen-Image, FLUX.2 dev and LTX-2.5 differ on selling outputs. What each page says, and hosted rows.
Written by Sume