Instagram posts in Google: draft the caption from the transcript
Public pro-account posts can appear in Google. Write the caption from what the video says; Sume video inspect returns a transcript at $0.01 per audio minute.

If a public post from a professional Instagram account can now show up in Google, the words you attach to it matter more than they did when it only lived in the app. The practical move is to write the caption from what the video actually says, and Sume's video inspect can give you that text: set transcribe: true and it returns the transcript for a clip you have imported, billed at $0.01 per audio minute (confirm in GET /v1/catalog). Sume does not publish to Instagram or control search indexing; this is only the drafting step.
Instagram's own help page on search engine indexing (read 2026-10-03) is where to check how public photos and videos are indexed and how to remove them from Google. Treat that page, not this one, as the authority on eligibility.
What Instagram says its search reads
Instagram's 2021 explanation of its own search says text matching is the most important signal, matched against usernames, bios, captions, hashtags and places, and it recommends keywords in captions rather than comments (read 2026-10-03). That page predates the Google change and describes Instagram's in-app search, so it is a hint about the value of caption keywords, not a statement about how Google ranks anything. What it does support is a simple habit: say the real topic in words, in the caption, near the start.
Pull the transcript
Video inspect needs a clip already on media.sume.com. Pass frames: false to skip stills, transcribe: true, an optional language_code hint, and segmentation.mode: "sentence" if you want caption-line-shaped segments. A silent clip fails with inspect_source_has_no_audio, so probe has_audio first if you are not sure. Default mode is sync, which waits up to 30 seconds and answers 200, or returns 202 to poll.
import json, os, urllib.request
API = "https://api.sume.com/v1"
def post(path, body, key):
req = urllib.request.Request(
f"{API}{path}",
data=json.dumps(body).encode(),
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": key,
},
method="POST",
)
with urllib.request.urlopen(req) as res:
return json.load(res)
res = post(
"/video-inspect",
{
"video_url": os.environ["SUME_CLIP_URL"],
"frames": False,
"transcribe": True,
"language_code": "en",
"segmentation": {"mode": "sentence"},
},
"inspect-transcript-001",
)
inspect = res.get("video_inspect", res)
transcript = inspect.get("transcript") or {}
text = transcript.get("text", "")
print(res["request_id"], len(text), "characters")
print(text[:160])
Turn it into a caption that can be found
Read the first lines of the transcript and look for the noun phrase a stranger would type: a product, a place, a how-to. Put that phrase in the first sentence of the caption in plain words, then add context. Use the transcript to catch words you said that the title card did not, such as a model number or a city.
A short checklist works better than a template:
- Is the main topic named in the first sentence, in the words a viewer would search?
- Are numbers, names and places spelled the way you said them on camera?
- Does the caption add something the video does not, such as a link target or a date?
- Did you read the public/professional account setting before assuming the post is eligible?
Limits to plan around
The transcript is for drafting. Review it before publishing: speech-to-text can misspell names, and a caption that repeats a mistake is worse than none. Whether Google surfaces a given post is Instagram's and Google's call, and Sume makes no claim about ranking.
| Item | Value (Sume docs) |
|---|---|
| Source length | Up to 1800 s per inspect |
| Transcript rate | $0.01 per audio minute |
| Duration hint | Omit and 1 minute is reserved; max hint 600 s |
| Language | language_code is a hint; omit for auto-detect |
| Silent clip | inspect_source_has_no_audio |
A worked example
Suppose a 40-second Reel is a founder explaining how to cut a hook with a trim tool. The title card says "Hook test". The transcript shows she actually said "trim the first three seconds" and named two clip lengths, 3 and 5 seconds. A caption that begins "How to trim the first 3 seconds of a Reel to test a hook" uses the words she spoke and the words someone would type. Compare that with a caption that only says "New tip!", which carries nothing a search could match.
Do the same for the second sentence: add the one fact the video does not say aloud, such as which tool you used or the date. Keep hashtags and place tags for what they are good at in the app, since Instagram's own search page lists hashtags and places among the fields its text matching reads.
What this does not do
It does not make a post eligible, and it does not make a private account indexable. Instagram's eligibility rules and the settings for turning indexing off live in its Help Center. Keep your own record of which posts were from professional accounts and public at publish time, so that if you change the setting you know what changed.
Sources
Related posts
More in Use cases
- Instagram Reel collab, up to 5 people: stitch their clips in one file
Meta says a Reel can have up to 5 collaborators. If each person films a clip, here is how to import them and stitch one 9:16 file with Sume Timeline.
- Does a TikTok watermark hurt Instagram Reels reach? What Meta says
Instagram's creator FAQ says a third-party watermark on your Reel may be limiting its reach. How to check a Reel's corners with Sume stills before you post.
- Instagram Reels Image Ad Specs: 1440x2560, 500 px Minimum
Meta lists Instagram Reels image ads as 9:16 at 1440x2560, 30 MB, 500 px minimum, 1% tolerance, 44 characters. Build the still with the Sume Images API.
- Reels audio types: original, licensed, voiceover, and an AI voice file
Instagram's Reels page names original audio, licensed music and Voiceover. Where a Sume-made voice file fits, and what the page does not say.
Written by Sume