Gemini CLI and the hosted Sume server: do not rely on env in headers
Add the hosted Sume server to Gemini CLI with httpUrl and a bearer header. Gemini expands env vars only in the env block; set a timeout above jobs_wait.

To use the hosted Sume server from Gemini CLI, add it under mcpServers in settings.json with httpUrl set to https://mcp.sume.com/mcp. If you authenticate with an API key, send Authorization: Bearer <key>, but know that the Gemini CLI docs describe environment-variable expansion only for the env block and say nothing about headers, so do not count on it there. Per the Gemini CLI MCP documentation (read 2026-10-06), the supported fields include httpUrl, headers, timeout, trust, includeTools and excludeTools.
That header rule is the trap. A config copied from another client with ${SUME_API_KEY} in a header may be sent as literal text if the client does not expand it, and the server then answers with an auth failure that looks like a bad key. Check what your Gemini CLI version actually sends before you trust the placeholder.
OAuth or an API key?
There are two ways to authenticate, per Sume hosted OAuth and API keys. OAuth gives a read-only session by default (mcp:read) and writes only if you switch Write on at consent. An API key, sent as a bearer header or x-api-key, gives the full hosted tool set, and writes and paid calls still need an idempotency_key. If the CLI cannot run the OAuth flow in your setup, the header route is the one to use, and then the key lives in a settings file.
How do I add it without a literal key in the file?
The command line avoids hand-editing JSON: the Gemini documentation shows gemini mcp add --transport http --header "Authorization: Bearer abc123" name URL. Your shell expands the variable before Gemini sees it, which is a legitimate way to get the value in, but the expanded key is then written into the settings file in plain text. Keep that file out of version control, prefer the user-level settings over a project file, and rotate the key if it ever lands in a repository. Remember that a literal key in any settings file is a secret at rest. A project-level settings file is the riskiest place for a literal key because it sits next to code that other people clone, so for a team, where several people share a repository, the better path is OAuth at the user level, with each developer signing in under their own account and a read-only grant until they need more.
What does the finished entry look like?
import json
entry = {
"mcpServers": {
"sume": {
"httpUrl": "https://mcp.sume.com/mcp",
"headers": {"Authorization": "Bearer REPLACE_AT_ADD_TIME"},
"timeout": 120000,
"excludeTools": [],
}
}
}
text = json.dumps(entry, indent=2)
cfg = json.loads(text)["mcpServers"]["sume"]
assert cfg["httpUrl"] == "https://mcp.sume.com/mcp"
assert "${" not in json.dumps(cfg["headers"]), "no env placeholder in headers"
assert cfg["timeout"] >= 60000, "jobs_wait can hold for up to 55 seconds"
print(text)Does the timeout setting matter for long jobs?
The timeout field matters because on the hosted server jobs_wait holds for 50 seconds by default and at most 55 (the API clamps larger values and reports wait_slice_clamped). The Gemini documentation lists a default of 600000 ms, so the default already covers a single wait, and the 120000 in the sketch is just a tighter value that still clears 55 seconds. The timeout is per request, so it limits one wait slice and not the whole job. If you set a short timeout, a healthy wait will be cut off and look like a failure; a transport-level 522 to 525 error is likewise not a job outcome, as covered in Jobs and results. When a wait slice ends, call jobs_wait again with the same ids and never resubmit the create.
After saving, start a session and ask the agent to call tools_list or mcp_health, the read-only checks the hosted quickstart suggests. Only then try a paid tool, with dry_run and max_spend_usd set. If you want the agent to stay read-only while you evaluate, the excludeTools list in the Gemini settings can hide tools you do not want it to see, which is a client-side guard that complements the server-side mcp:read scope. In practice, start with the smallest useful set: the read tools for discovery, jobs_wait and jobs_result for results, and one creation tool you actually plan to use. Check the exact tool names with tools_list first, rather than guessing them, since the live ids use underscores.
Sources
Related posts
More in Developers
- Get the video URL from a Sume webhook: pick the artifact by type
A Sume job.completed payload lists artifacts with id, url, type and content_type. Select the video by content_type, not array index. TypeScript for Node.
- Go net/http client for Sume images: handle 200 and 202 on a model swap
A Go program that posts to Sume /v1/images with the model id from an env var and branches on 200 versus 202, so a gpt-image-1 swap is not a rebuild.
- gpt-image-1 returns 404 model_not_found on Sume: which id to send
gpt-image-1, gpt-image-1.5 or gemini-2.5-flash-image sent to Sume /v1/images return 404 model_not_found. Ids to send instead, plus a lookup.
- Heroku H12 at 30 seconds: call Sume async, not sync
Heroku's router ends a request at 30 s (H12). Sume sync mode waits up to 30 s too. Submit async, return 202 to the browser, then poll or take a webhook.
Written by Sume