AuK: MIT speech model that edits audio, and what Sume covers
AuK is Tencent's 1.5B MIT-licensed speech model for generation and editing. What its card lists, what it omits, and which parts Sume's audio tools cover.

AuK is a 1.5B-parameter speech model from Tencent, released under the MIT License, that generates speech and also edits it from a plain-language instruction: rewrite the words, change pitch, speed or volume, shift emotion, or clean the take up. If you only need new narration, a hosted text-to-speech job is the shorter path. If you need to repair speech that already exists, AuK is the open-weights option worth testing, and Sume does not offer that repair step.
This post separates what the AuK model cards say from what Sume's docs cover. Everything about AuK comes from the AuK card and the AuK-Flash card, both read on 2026-10-03.
What does AuK actually do?
The card describes AuK as a foundation model for speech generation and editing, trained on millions of hours of audio, driven by a single natural-language instruction interface. A multimodal language model encoder (Qwen2.5-Omni-3B) reads the instruction and the audio. The card groups the tasks into five families.
Two checkpoints are listed: AuK, the base model, and AuK-Flash, a distilled variant the card describes as 4-step inference. The card gives no sample rate, maximum clip length, language list or GPU memory figure, so treat all four as unknown until you run it yourself.
| Family | What the card lists | Closest Sume surface |
|---|---|---|
| Generation | Zero-shot and instruction-based text-to-speech | tts_create over hosted MCP |
| Content editing | Speech and lyric rewriting | None; regenerate the line |
| Acoustic editing | Pitch, speed and volume | None |
| Paralinguistic editing | Emotion, timbre, de-accent, whisper conversion | None |
| Enhancement and separation | Denoising, dereverberation, speaker extraction | None (audio detach only extracts a track) |
What does Sume cover for speech and audio?
Sume's MCP tool list names tts_create as a paid create tool and tts_source_get and tts_source_verify_spine as free reads that check selected TTS jobs against an accepted script. The basics page says TTS also exists as an HTTP API. The docs I read do not describe a language count, voice cloning field or instruction-based editing for TTS, so do not plan around those.
Around the speech itself, the audio tools are mechanical. Audio detach pulls the audio track of one Sume-hosted video into a wav or mp3 for $0.01 per job, with a source up to 1800 seconds and an output up to 900 seconds. Timeline audio concatenates up to 20 parts or splits a file by range in the sample domain, with no re-synthesis. Video inspect can transcribe at $0.01 per audio minute.
When is AuK the better tool, and when is a hosted job?
Pick by whether the audio already exists.
The honest gap: if a finished take has one wrong word, Sume's route is to regenerate that sentence and splice it with timeline audio. That works for clean narration and costs one more TTS job. It will not fix room echo or change the emotion of a recording you cannot redo.
- Existing take needs a fix (swap a word, remove reverb, extract one speaker): AuK or a similar editor, run on your own GPU.
- New narration at volume with no GPU to run: a hosted TTS job, then detach, split and concat to assemble it.
- Commercial use: the card states MIT for AuK, which is permissive. The card does not cover training data or voice-rights questions, so clear consent for any real voice yourself.
- Anything you need to pin to a language, sample rate or latency: test on your hardware first, since the card publishes none of those.
What should you test before relying on AuK?
Run three checks on your own clips. First, edit one word in a 20-second take and listen at the splice; speech editors can drift in timbre across the edit. Second, try AuK-Flash against the base checkpoint on the same instruction, because a 4-step model trades something for speed and the card does not say what. Third, measure memory on your GPU, since none is published.
If the edit holds up, keep AuK as a repair step before assembly. Detach or import the repaired audio, then use timeline audio to join it with the other takes. If it does not, regenerating the line through a hosted job is usually the cheaper failure.
Sources
Related posts
More in Models
- Can you sell images from open-weights models? Licences compared
Open weights do not mean commercial use. Ideogram 4, Qwen-Image, FLUX.2 dev and LTX-2.5 differ on selling outputs. What each page says, and hosted rows.
- ChatGPT Try On from a screenshot: the same edit through an API
ChatGPT Try On starts from a selfie plus a product screenshot. Do the same edit with openai/gpt-image-2.5 on Sume: two references, one prompt, one Python call.
- Chinese, Hindi and Russian text to speech API: Sume zh, hi and ru
Sume's Voices library has zh, hi and ru tags. How to request each, what Eleven v4 lists, and the one field that stops an English-sounding read.
- Clone a voice from an MP4 or WebM file: Sume Voices takes video, 20 MB
Sume's Voices clone accepts wav, mp3, m4a, mp4, webm and ogg up to 20 MB. How to prepare a clean clip from a video, and what the upload check rejects.
Written by Sume