Danish, Norwegian, Finnish text to speech API: da, no, fi on Sume
Sonic 3.6 lists da, no and fi. Norwegian is the code no, not nb or nn. How to send each through Sume TTS, which voice id works, and what a script costs.

Nordic narration on Sume TTS means setting language to da for Danish, no for Norwegian or fi for Finnish. Cartesia's Sonic 3.6 page lists all three, with Swedish (sv) alongside; the same three codes appear on Sonic 3.5 and Sonic 3. One trap: Cartesia's list has a single Norwegian entry, no, so if your content system tags text as nb (Bokmal) or nn (Nynorsk), map it to no before it reaches the request.
Swedish is the only Nordic language in Sume's voice-library list. Danish, Norwegian and Finnish can be spoken, but a voice for them is a UUID you bring, not a row you create in the library.
Which codes does Sume forward?
Sume's contract accepts any language string from 2 to 16 characters, described as BCP-47 or ISO 639, and forwards it. There is no allowlist in the Sume worker for it, which means an unsupported or misspelled code reaches Cartesia instead of failing at the Sume edge. Check the code against the table before you template it into a pipeline.
| Language | Code | sonic-3.6-2026-08-27 | sonic-3.5-2026-05-04 | sonic-3-2026-01-12 | In Sume's voice-library language list |
|---|---|---|---|---|---|
| Danish | da | Yes | Yes | Yes | No |
| Norwegian | no | Yes | Yes | Yes | No |
| Finnish | fi | Yes | Yes | Yes | No |
| Swedish | sv | Yes | Yes | Yes | Yes |
Is the voice language checked?
Only for voices Sume knows. If the voice.id matches a voice saved in your Sume library, its primary language is compared with the request, and a known mismatch is a 409 tts_voice_language_mismatch with no job and no charge. Regional tags compare by primary language, so da-DK against a Danish voice passes. A raw UUID with no library row is not blocked, because the library never gates a raw voice id.
In the hosted MCP the same mismatch arrives as a non-error warning that needs the user's confirmation, after which the identical request is retried with confirm_language_mismatch: true.
What does a Danish, Norwegian or Finnish script cost?
One price book covers every language: $0.0475 per 1,000 transcript characters. Finnish and Danish run longer than English per idea, so budget by characters. A 6,000-character Finnish script is $0.285; the cap is 20,000 characters or 1,200 seconds of audio per job.
- Defaults: MP3, 44,100 Hz, 128 kbps. Pick
wavif you plan to cut or join takes. generation_config.speedruns 0.6 to 1.5 andvolume0.5 to 2, for matching a ten-second slot.- Ask for
timestamps.wordswhen you will caption the result.
What does the request look like?
Example for Norwegian, with VOICE_ID set to a Norwegian-capable voice UUID:
import os
import uuid
import requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={
"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"transcript": "Velkommen til ukens oppdatering.",
"language": "no",
"voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
"mode": "async",
},
timeout=30,
)
r.raise_for_status()
print(r.json())What should I test before a release?
Run one paragraph per language and listen to it. Check three things: that numbers and dates read the way your market writes them, that a product name is not read as a native word, and that the voice sounds like the language and not like an English voice reading foreign text. A short sample of 200 characters costs $0.0095 at Sume's rate, which is cheap compared with re-recording a finished video. Write the transcript exactly as it should be spoken; Sume has no transcript-rewriting step.
Sources
Related posts
More in Models
- DeepSeek V4 Pro API after Sept 14: is there a Sume row?
DeepSeek says V4 Pro API service continues past Sept 14 at unchanged billing. Sume's agent catalog lists V4.1 Flash and older Flash rows, but no V4 Pro.
- Did OpenAI change GPT Image 2.5 since launch? Changelog check, Oct 3
OpenAI's changelog shows GPT Image 2.5 shipped Sept 8 and no October image entry. The Sept 25 image fix covers GPT-6 inputs, not generation.
- DMAD 4-step MiniMax H3: 50-step baseline vs Turbo LoRA
DMAD distills MiniMax H3 from 50 steps to 4. Why that 4 is not the Turbo LoRA's 4, what the card says, what it omits, and where hosted jobs fit.
- Edit a video with a prompt and a photo: Omni edit takes no references
Gemini Omni Flash 1.1's video edit on Sume takes a video_url and a prompt only. To put a specific person from a photo into a clip, use h3-max-recast instead.
Written by Sume