AuK: MIT speech model that edits audio, and what Sume covers

AuK is Tencent's 1.5B MIT-licensed speech model for generation and editing. What its card lists, what it omits, and which parts Sume's audio tools cover.

5 min readSume
All posts

AuK is a 1.5B-parameter speech model from Tencent, released under the MIT License, that generates speech and also edits it from a plain-language instruction: rewrite the words, change pitch, speed or volume, shift emotion, or clean the take up. If you only need new narration, a hosted text-to-speech job is the shorter path. If you need to repair speech that already exists, AuK is the open-weights option worth testing, and Sume does not offer that repair step.

This post separates what the AuK model cards say from what Sume's docs cover. Everything about AuK comes from the AuK card and the AuK-Flash card, both read on 2026-10-03.

What does AuK actually do?

The card describes AuK as a foundation model for speech generation and editing, trained on millions of hours of audio, driven by a single natural-language instruction interface. A multimodal language model encoder (Qwen2.5-Omni-3B) reads the instruction and the audio. The card groups the tasks into five families.

Two checkpoints are listed: AuK, the base model, and AuK-Flash, a distilled variant the card describes as 4-step inference. The card gives no sample rate, maximum clip length, language list or GPU memory figure, so treat all four as unknown until you run it yourself.

Task families listed on the AuK card, read 2026-10-03.
FamilyWhat the card listsClosest Sume surface
GenerationZero-shot and instruction-based text-to-speechtts_create over hosted MCP
Content editingSpeech and lyric rewritingNone; regenerate the line
Acoustic editingPitch, speed and volumeNone
Paralinguistic editingEmotion, timbre, de-accent, whisper conversionNone
Enhancement and separationDenoising, dereverberation, speaker extractionNone (audio detach only extracts a track)

What does Sume cover for speech and audio?

Sume's MCP tool list names tts_create as a paid create tool and tts_source_get and tts_source_verify_spine as free reads that check selected TTS jobs against an accepted script. The basics page says TTS also exists as an HTTP API. The docs I read do not describe a language count, voice cloning field or instruction-based editing for TTS, so do not plan around those.

Around the speech itself, the audio tools are mechanical. Audio detach pulls the audio track of one Sume-hosted video into a wav or mp3 for $0.01 per job, with a source up to 1800 seconds and an output up to 900 seconds. Timeline audio concatenates up to 20 parts or splits a file by range in the sample domain, with no re-synthesis. Video inspect can transcribe at $0.01 per audio minute.

When is AuK the better tool, and when is a hosted job?

Pick by whether the audio already exists.

The honest gap: if a finished take has one wrong word, Sume's route is to regenerate that sentence and splice it with timeline audio. That works for clean narration and costs one more TTS job. It will not fix room echo or change the emotion of a recording you cannot redo.

  • Existing take needs a fix (swap a word, remove reverb, extract one speaker): AuK or a similar editor, run on your own GPU.
  • New narration at volume with no GPU to run: a hosted TTS job, then detach, split and concat to assemble it.
  • Commercial use: the card states MIT for AuK, which is permissive. The card does not cover training data or voice-rights questions, so clear consent for any real voice yourself.
  • Anything you need to pin to a language, sample rate or latency: test on your hardware first, since the card publishes none of those.

What should you test before relying on AuK?

Run three checks on your own clips. First, edit one word in a 20-second take and listen at the splice; speech editors can drift in timbre across the edit. Second, try AuK-Flash against the base checkpoint on the same instruction, because a 4-step model trades something for speed and the card does not say what. Third, measure memory on your GPU, since none is published.

If the edit holds up, keep AuK as a repair step before assembly. Detach or import the repaired audio, then use timeline audio to join it with the other takes. If it does not, regenerating the line through a hosted job is usually the cheaper failure.

Sources

Related posts

More in Models

All Models posts

Written by Sume