EmbeddingGemma 2 audio embeddings or a transcript: what to index
EmbeddingGemma 2 embeds audio directly. Sume has no embedding endpoint, but STT gives text and word timings at $0.01 per minute for search.

If you want to search a pile of spoken clips, there are now two routes: embed the audio itself, or transcribe it and index the text. EmbeddingGemma 2, reported in the October 7 AI brief (read 2026-10-07), is a 740M multimodal embedding model under Apache 2.0 that covers text, code, images, video and audio. That is the first route.
Sume does not ship an embedding model. The public model catalog and the API reference list generators and media tools, and none of them returns vectors, so the embedding step stays on your side. What Sume does give you for the second route is speech-to-text: Sume STT 1.0 returns the transcript plus word timings, at $0.01 per audio minute (read 2026-10-07 from the repo's catalog rate card).
Which route fits which job
Embedding the audio helps when the sound itself matters: a laugh, a jingle, a room tone. Indexing a transcript helps when people search by what was said, and it gives you a timestamp to jump to. A transcript is also something you can read, grep and fix.
| Question | Embed the audio | Transcribe, then index text |
|---|---|---|
| Find the clip where someone says a product name | Weak: depends on the model | Strong: exact words |
| Jump to the second it was said | You must add your own chunking | Word timings come back with the text |
| Find clips that sound like a given jingle | Strong | Not possible |
| Cost on Sume | Not offered by Sume | $0.01 per audio minute, up to 10 minutes per job |
The Sume side: one STT call per clip
POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url (a Sume media URL is preferred), an optional language_code hint such as ko, and an optional duration_seconds between 1 and 600 that makes the cost reservation exact. Without the hint Sume reserves one minute. One job covers at most 10 minutes, so a 25-minute recording is three jobs, which you can cut with audio detach and its range field.
The docs note that some generators, STT among them, can appear on the dev API before the production OpenAPI snapshot lists them, so read GET /v1/catalog with your key to see what your account can call.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: index-clip-001" \
-d '{"audio_url": "https://media.sume.com/artifacts/artf_demo/call.wav", "duration_seconds": 90}'What it costs to index a library
Transcription is cheap enough that the embedding model, not the transcript, is the cost you will tune. At the fixed rate, 100 hours of clips is 6,000 minutes, or $60.00 before any retries; the same hour split into 10-minute jobs rounds each job up to the next cent, so a job of 90 seconds is $0.02. Embed the transcript with EmbeddingGemma 2 or any text model you already use, and keep the word timings next to each chunk so a search hit can open the clip at the right second.
A quick rule
Start with the transcript. It answers most spoken-clip searches, it costs a cent a minute on Sume, and it leaves the door open: if you later add audio embeddings for sound-alike search, the transcript index stays useful beside it.
Sources
Related posts
More in Models
- EmbeddingGemma 2 for video search: Sume has no embeddings endpoint
EmbeddingGemma 2 is reported as a 740M, Apache 2.0 multimodal embedder. Sume sells no embeddings API as of 2026-10-07; here is what it does sell for search.
- Fix extra fingers in an AI image: a masked GPT Image 2.5 edit
Fix a bad hand in an AI image with a mask_url edit on openai/gpt-image-2.5: mask only the hand, keep the rest, and budget about $0.0835 per try at high.
- FLUX 3 Image 5456x3072 at 300 dpi: print size vs Sume pixel caps
BFL shows FLUX 3 Image at 5456x3072 (16.8 MP). At 300 dpi that prints 18.19 x 10.24 in. Here is what Sume models give at the same dpi.
- gemini-2.5-flash-image: Google pages give Oct 2, 2026 and Mar 15, 2027
Google's pricing page says gemini-2.5-flash-image shuts down Oct 2, 2026; its deprecations page says Mar 15, 2027. How to plan around two dates.
Written by Sume