EmbeddingGemma 2 audio embeddings or a transcript: what to index

EmbeddingGemma 2 embeds audio directly. Sume has no embedding endpoint, but STT gives text and word timings at $0.01 per minute for search.

5 min readSume
All posts

If you want to search a pile of spoken clips, there are now two routes: embed the audio itself, or transcribe it and index the text. EmbeddingGemma 2, reported in the October 7 AI brief (read 2026-10-07), is a 740M multimodal embedding model under Apache 2.0 that covers text, code, images, video and audio. That is the first route.

Sume does not ship an embedding model. The public model catalog and the API reference list generators and media tools, and none of them returns vectors, so the embedding step stays on your side. What Sume does give you for the second route is speech-to-text: Sume STT 1.0 returns the transcript plus word timings, at $0.01 per audio minute (read 2026-10-07 from the repo's catalog rate card).

Which route fits which job

Embedding the audio helps when the sound itself matters: a laugh, a jingle, a room tone. Indexing a transcript helps when people search by what was said, and it gives you a timestamp to jump to. A transcript is also something you can read, grep and fix.

Choosing an index for spoken clips, read 2026-10-07
QuestionEmbed the audioTranscribe, then index text
Find the clip where someone says a product nameWeak: depends on the modelStrong: exact words
Jump to the second it was saidYou must add your own chunkingWord timings come back with the text
Find clips that sound like a given jingleStrongNot possible
Cost on SumeNot offered by Sume$0.01 per audio minute, up to 10 minutes per job

The Sume side: one STT call per clip

POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url (a Sume media URL is preferred), an optional language_code hint such as ko, and an optional duration_seconds between 1 and 600 that makes the cost reservation exact. Without the hint Sume reserves one minute. One job covers at most 10 minutes, so a 25-minute recording is three jobs, which you can cut with audio detach and its range field.

The docs note that some generators, STT among them, can appear on the dev API before the production OpenAPI snapshot lists them, so read GET /v1/catalog with your key to see what your account can call.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: index-clip-001" \
  -d '{"audio_url": "https://media.sume.com/artifacts/artf_demo/call.wav", "duration_seconds": 90}'

What it costs to index a library

Transcription is cheap enough that the embedding model, not the transcript, is the cost you will tune. At the fixed rate, 100 hours of clips is 6,000 minutes, or $60.00 before any retries; the same hour split into 10-minute jobs rounds each job up to the next cent, so a job of 90 seconds is $0.02. Embed the transcript with EmbeddingGemma 2 or any text model you already use, and keep the word timings next to each chunk so a search hit can open the clip at the right second.

A quick rule

Start with the transcript. It answers most spoken-clip searches, it costs a cent a minute on Sume, and it leaves the door open: if you later add audio embeddings for sound-alike search, the transcript index stays useful beside it.

Sources

Related posts

More in Models

All Models posts

Written by Sume