Danish, Norwegian, Finnish text to speech API: da, no, fi on Sume

Sonic 3.6 lists da, no and fi. Norwegian is the code no, not nb or nn. How to send each through Sume TTS, which voice id works, and what a script costs.

4 min readSume
All posts

Nordic narration on Sume TTS means setting language to da for Danish, no for Norwegian or fi for Finnish. Cartesia's Sonic 3.6 page lists all three, with Swedish (sv) alongside; the same three codes appear on Sonic 3.5 and Sonic 3. One trap: Cartesia's list has a single Norwegian entry, no, so if your content system tags text as nb (Bokmal) or nn (Nynorsk), map it to no before it reaches the request.

Swedish is the only Nordic language in Sume's voice-library list. Danish, Norwegian and Finnish can be spoken, but a voice for them is a UUID you bring, not a row you create in the library.

Which codes does Sume forward?

Sume's contract accepts any language string from 2 to 16 characters, described as BCP-47 or ISO 639, and forwards it. There is no allowlist in the Sume worker for it, which means an unsupported or misspelled code reaches Cartesia instead of failing at the Sume edge. Check the code against the table before you template it into a pipeline.

Language codes per Cartesia's model pages (read 2026-10-03); the last column is Sume's voice-library list from its voice code
LanguageCodesonic-3.6-2026-08-27sonic-3.5-2026-05-04sonic-3-2026-01-12In Sume's voice-library language list
DanishdaYesYesYesNo
NorwegiannoYesYesYesNo
FinnishfiYesYesYesNo
SwedishsvYesYesYesYes

Is the voice language checked?

Only for voices Sume knows. If the voice.id matches a voice saved in your Sume library, its primary language is compared with the request, and a known mismatch is a 409 tts_voice_language_mismatch with no job and no charge. Regional tags compare by primary language, so da-DK against a Danish voice passes. A raw UUID with no library row is not blocked, because the library never gates a raw voice id.

In the hosted MCP the same mismatch arrives as a non-error warning that needs the user's confirmation, after which the identical request is retried with confirm_language_mismatch: true.

What does a Danish, Norwegian or Finnish script cost?

One price book covers every language: $0.0475 per 1,000 transcript characters. Finnish and Danish run longer than English per idea, so budget by characters. A 6,000-character Finnish script is $0.285; the cap is 20,000 characters or 1,200 seconds of audio per job.

  • Defaults: MP3, 44,100 Hz, 128 kbps. Pick wav if you plan to cut or join takes.
  • generation_config.speed runs 0.6 to 1.5 and volume 0.5 to 2, for matching a ten-second slot.
  • Ask for timestamps.words when you will caption the result.

What does the request look like?

Example for Norwegian, with VOICE_ID set to a Norwegian-capable voice UUID:

import os
import uuid

import requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={
        "x-api-key": os.environ["SUME_API_KEY"],
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "transcript": "Velkommen til ukens oppdatering.",
        "language": "no",
        "voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
        "mode": "async",
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

What should I test before a release?

Run one paragraph per language and listen to it. Check three things: that numbers and dates read the way your market writes them, that a product name is not read as a native word, and that the voice sounds like the language and not like an English voice reading foreign text. A short sample of 200 characters costs $0.0095 at Sume's rate, which is cheap compared with re-recording a finished video. Write the transcript exactly as it should be spoken; Sume has no transcript-rewriting step.

Sources

Related posts

More in Models

All Models posts

Written by Sume